Pith. sign in

REVIEW 2 major objections 1 minor 1 cited by

Accommodation Goes Both Ways: Studying Linguistic Convergence Between Humans and Language Models

T0 review · 2 major / 1 minor · reviewed 2026-06-29 · grok-4.3

Pith's one-line read LLMs overconverge to users' linguistic style far more than humans accommodate back.

desk verdict LLMs overconverge to user style in WildChat while humans match human-human baselines, but the metric may not separate accommodation from model conditioning or corpus effects. read the letter →

arxiv 2605.29278 v1 pith:BDDP2NYH submitted 2026-05-28 cs.CL

classification cs.CL
keywords linguisticconvergenceaccommodationhuman-LLMdialogueLLMbehaviorWildChatcorpusasymmetricfunctionwordsopen-classfeatures
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper measures how humans and LLMs shift their language toward each other in real multi-turn chats. It applies an asymmetric convergence metric to the WildChat corpus of ChatGPT transcripts across eight languages. LLMs adapt strongly on both function words and open-class vocabulary. Human adaptation rates stay in line with ordinary human-to-human conversation. The result points to one-sided accommodation in human-LLM dialogue.

What carries the argument

The asymmetric convergence metric, which separately quantifies each party's adaptation to the other's linguistic features on function words and open-class items in multi-turn WildChat transcripts.

What would settle it

A new corpus of human-LLM conversations, processed with the identical asymmetric metric, that shows human convergence rates significantly above or below human-human baselines would falsify the central claim.

Watch

Extended reading notes

Core claim

Using an asymmetric convergence metric on WildChat, the study shows that LLMs significantly overconverge toward their users on both function word and open-class features across eight languages, while human convergence rates remain consistent with human-human baselines, indicating asymmetric accommodation in human-LLM dialogue.

Load-bearing premise

The asymmetric convergence metric applied to the WildChat corpus isolates genuine linguistic accommodation rather than artifacts of the data collection process or the specific LLM behavior.

Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 1 minor

Summary. The paper claims that LLMs significantly overconverge toward users on function-word and open-class features across eight languages in WildChat transcripts, while human convergence rates remain consistent with human-human baselines, indicating asymmetric accommodation where LLMs overfit to user style but humans accommodate LLMs as they would other people.

Significance. If the asymmetric metric robustly isolates accommodation, the cross-linguistic scale and real-world corpus would provide valuable evidence on how LLMs shape linguistic behavior differently from human interlocutors, with implications for conversational AI design and theories of accommodation.

major comments (2)
  1. [Methods] Methods (asymmetric convergence metric): the abstract and central claim rely on this metric isolating genuine accommodation, yet no equation, normalization procedure, or control condition is described to rule out LLM conditioning on prior turns (which induces overlap by construction) or WildChat collection/filtering artifacts; without these, the overconvergence finding and human-baseline comparison are uninterpretable.
  2. [Results] Results (human-human baseline comparison): the claim that human rates are 'broadly consistent' with prior baselines requires explicit statistical controls, exclusion criteria, and source details for the baselines; the abstract provides none, leaving open whether differences in dialogue length, topic, or participant pool confound the consistency conclusion.
minor comments (1)
  1. [Abstract] Abstract: metric definition, baseline sources, and statistical controls should be summarized briefly to allow readers to assess the findings without the full text.

Simulated Author's Rebuttal

2 responses · 0 unresolved

We thank the referee for their constructive comments. We address each major point below and will incorporate clarifications in the revised manuscript.

read point-by-point responses
  1. Referee: [Methods] Methods (asymmetric convergence metric): the abstract and central claim rely on this metric isolating genuine accommodation, yet no equation, normalization procedure, or control condition is described to rule out LLM conditioning on prior turns (which induces overlap by construction) or WildChat collection/filtering artifacts; without these, the overconvergence finding and human-baseline comparison are uninterpretable.

    Authors: We agree that the manuscript would be strengthened by an explicit presentation of the asymmetric convergence metric. In the revision we will add the full equation, describe the normalization steps, and include control analyses that isolate accommodation from simple prior-turn conditioning and from WildChat filtering effects. These controls will be reported alongside the main results. revision: yes

  2. Referee: [Results] Results (human-human baseline comparison): the claim that human rates are 'broadly consistent' with prior baselines requires explicit statistical controls, exclusion criteria, and source details for the baselines; the abstract provides none, leaving open whether differences in dialogue length, topic, or participant pool confound the consistency conclusion.

    Authors: We will expand the results section to cite the exact baseline sources, report statistical comparisons that control for dialogue length and topic, and state the exclusion criteria used. This will allow readers to evaluate whether the observed consistency holds after accounting for potential confounds. revision: yes

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity; empirical analysis relies on external corpus and baselines

full rationale

The paper is an empirical study applying an asymmetric convergence metric to the external WildChat corpus and comparing human rates to prior human-human baselines. No equations, derivations, fitted parameters presented as predictions, or self-citation chains that reduce the central claims to inputs by construction appear in the abstract or described structure. The load-bearing steps are data-driven measurements against independent references, satisfying the criteria for a self-contained result with no circular reduction.

Assumptions & free parameters 0 free parameters · 1 assumptions · 0 invented entities

Central claim depends on validity of the convergence metric and representativeness of WildChat for natural human-LLM dialogue; no free parameters or invented entities are visible in the abstract.

assumptions (1)
  • domain assumption The asymmetric convergence metric accurately measures linguistic accommodation independent of corpus-specific artifacts.
    Metric is the load-bearing tool used to compare LLM and human rates.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Accommodation Goes Both Ways: Studying Linguistic Convergence Between Humans and Language Models." pith.science (2026). https://pith.science/paper/BDDP2NYH

@misc{pith2026260529278,
  author       = {Pith},
  title        = {Pith review of: Accommodation Goes Both Ways: Studying Linguistic Convergence Between Humans and Language Models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/BDDP2NYH}},
  note         = {Machine review of arXiv:2605.29278}
}
read the original abstract

As LLMs become increasingly integrated into daily life, understanding how their presence will shape human linguistic behavior is an open question. We present a large-scale study of linguistic convergence in human-LLM dialogue, examining how humans and LLMs accommodate each other's linguistic style during multi-turn conversations. Using an asymmetric convergence metric on WildChat, a corpus of real-world ChatGPT transcripts, we find that while LLMs significantly overconverge toward their users on both function word and open-class features across eight languages, human convergence rates in this setting are broadly consistent with human-human baselines. These findings suggest that accommodation in human-LLM dialogue is asymmetric: while LLMs dramatically overfit to their users' style, humans linguistically accommodate LLMs no differently than they would another person.

Figures

Figures reproduced from arXiv: 2605.29278 by the authors.

Figure 1
Figure 1. Mean (a) LIWC and (b) NOUN scores for LLMs and users across eight languages in WildChat. LLMs consistently converge more than users across all languages; the reported values indicate ∆ = LLM - User. for LLM training, falls in the middle of this range rather than at the top or bottom for any individ￾ual feature, suggesting the asymmetry is not sim￾ply a function of model quality on a given lan￾guage. These same patte… view at source ↗
Figure 2
Figure 2. Mean LIWC convergence scores per turn position for LLMs and users across eight WildChat languages. [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Mean NOUN convergence scores per turn position for LLMs and users across eight WildChat languages. [PITH_FULL_IMAGE:figures/full_fig_p009_3.png] view at source ↗

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Instruction-Tuned Models Locally Reuse Human Syntax More Than Humans Do

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Instruction-tuned LLMs reuse the immediate human turn's grammar more than the matched human response does in raw output, but the effect reverses when response structure is held constant.

Reference graph

Works this paper leans on

35 extracted references · 4 canonical work pages · cited by 1 Pith paper

  1. [1]

    Alberto Agosti and Alessandra Rellini. 2007. The italian liwc dictionary. Austin, TX: LIWC. Net

  2. [2]

    Pedro Balage Filho, Thiago Alexandre Salgueiro Pardo, and Sandra Alu \' sio. 2013. An evaluation of the brazilian portuguese liwc dictionary for sentiment analysis. In Proceedings of the 9th Brazilian Symposium in Information and Human Language Technology

  3. [3]

    Anshul Bawa, Monojit Choudhury, and Kalika Bali. 2018. Accommodation of conversational code-choice. In Proceedings of the Third Workshop on Computational Approaches to Linguistic Code-Switching, pages 82--91

  4. [4]

    Aleksandrs Berdi c evskis and Viktor Erbro. 2023. You say tomato, i say the same: A large-scale study of linguistic accommodation in online communities. In Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa), pages 415--424

  5. [5]

    Paras Bhatt and Anthony Rios. 2021. Detecting bot-generated text by characterizing linguistic accommodation in human-bot interactions. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 3235--3247

  6. [6]

    Terra Blevins, Susanne Schmalwieser, and Benjamin Roth. 2026. https://doi.org/10.18653/v1/2026.eacl-long.34 Do language models accommodate their users? a study of linguistic convergence . In Proceedings of the 19th Conference of the E uropean Chapter of the A ssociation for C omputational L inguistics (Volume 1: Long Papers) , pages 791--807, Rabat, Moroc...

  7. [7]

    Francis Bond and Ryan Foster. 2013. Linking and extending an open multilingual wordnet. In Proceedings of the 51st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1352--1362

  8. [8]

    Ryan L Boyd, Ashwini Ashokkumar, Sarah Seraj, and James W Pennebaker. 2022. The development and psychometric properties of liwc-22. Austin, TX: University of Texas at Austin, 10(1-47):6

Show all 35 references
  1. [9]

    Susan E Brennan. 1996. Lexical entrainment in spontaneous dialog. Proceedings of ISSD, 96:41--44

  2. [10]

    Susan E Brennan and Herbert H Clark. 1996. Conceptual pacts and lexical choice in conversation. Journal of experimental psychology: Learning, memory, and cognition, 22(6):1482

  3. [11]

    Pengbo Chen, Huining Guan, and Eui Jun Jeong. 2026. Who accommodates whom? bidirectional linguistic accommodation and progressive interpersonal convergence in human--ai conversations. Behavioral Sciences

  4. [12]

    Cindy K Chung and James W Pennebaker. 2012. Linguistic inquiry and word count (liwc): pronounced “luke,”... and other useful facts. In Applied natural language processing: Identification, investigation and resolution, pages 206--229. IGI Global Scientific Publishing

  5. [13]

    Cristian Danescu-Niculescu-Mizil and Lillian Lee. 2011. Chameleons in imagined conversations: A new approach to understanding coordination of linguistic style in dialogs. In Proceedings of the 2nd workshop on cognitive modeling and computational linguistics, pages 76--87

  6. [14]

    Christiane Fellbaum. 1998. WordNet: An electronic lexical database. MIT press

  7. [15]

    Howard Giles, Nikolas Coupland, and Justine Coupland. 1991. Accommodation theory: Communication, context, and consequence. Contexts of accommodation: Developments in applied sociolinguistics, 1:1--68

  8. [16]

    Matthew Honnibal, Ines Montani, Sofie Van Landeghem, Adriane Boyd, and 1 others. 2020. spacy: Industrial-strength natural language processing in python

  9. [17]

    Molly E Ireland, Richard B Slatcher, Paul W Eastwick, Lauren E Scissors, Eli J Finkel, and James W Pennebaker. 2011. Language style matching predicts relationship initiation and stability. Psychological science, 22(1):39--44

  10. [18]

    Cameron R Jones and Benjamin K Bergen. 2025. Large language models pass the turing test. arXiv preprint arXiv:2503.23674

  11. [19]

    Andreas Kailer and Cindy K Chung. 2011. The russian liwc2007 dictionary. Austin, TX: LIWC. net

  12. [20]

    Florian Kandra, Vera Demberg, and Alexander Koller. 2025. https://doi.org/10.18653/v1/2025.acl-short.68 LLM s syntactically adapt their language use to their conversational partner . In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Vo...

  13. [21]

    Yanran Li, Hui Su, Xiaoyu Shen, Wenjie Li, Ziqiang Cao, and Shuzi Niu. 2017. Dailydialog: A manually labelled multi-turn dialogue dataset. In Proceedings of the Eighth International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 986--995

  14. [22]

    Ryan Lowe, Nissan Pow, Iulian Vlad Serban, and Joelle Pineau. 2015. The U buntu D ialogue C orpus: A large dataset for research in unstructured multi-turn dialogue systems. In Proceedings of the 16th annual meeting of the special interest group on discourse and dialogue, pages...

  15. [23]

    Yuwen Lyu, Julian Chun-Chung Chow, Ji-Jen Hwang, Zhi Li, Cheng Ren, and Jungui Xie. 2022. Psychological well-being of left-behind children in china: text mining of the social media website zhihu. International journal of environmental research and public health, 19(4):2127

  16. [24]

    Arjun Mukherjee and Bing Liu. 2012. Analysis of linguistic style accommodation in online debates. In Proceedings of COLING 2012, pages 1831--1846

  17. [25]

    Clifford Nass and Youngme Moon. 2000. Machines and mindlessness: Social responses to computers. Journal of social issues, 56(1):81--103

  18. [26]

    Clifford Nass, Jonathan Steuer, and Ellen R Tauber. 1994. Computers are social actors. In Proceedings of the SIGCHI conference on Human factors in computing systems, pages 72--78

  19. [27]

    Kate G Niederhoffer and James W Pennebaker. 2002. Linguistic style matching in social interaction. Journal of language and social psychology, 21(4):337--360

  20. [28]

    Martin J Pickering and Simon Garrod. 2004. Toward a mechanistic psychology of dialogue. Behavioral and brain sciences, 27(2):169--190

  21. [29]

    Annie Piolat, RJ Booth, Cindy K Chung, M Davids, and JW Pennebaker. 2011. The french dictionary for liwc: Modalities of construction and examples of use. Psychologie fran c aise , 56(3):145--159

  22. [30]

    Nair \'a n Ram \' rez-Esparza, James W Pennebaker, Florencia Andrea Garc \' a, and Raquel Suri \'a . 2007. La psicolog \' a del uso de las palabras: Un programa de computadora que analiza textos en espa \ n ol. Revista mexicana de psicolog \' a , 24(1):85--99

  23. [31]

    Arthur Ward and Diane J Litman. 2007. Automatically measuring lexical and acoustic/prosodic convergence in tutorial dialog corpora. In SLaTE, pages 57--60

  24. [32]

    Fulei Zhang and Zhou Yu. 2025. Mind the gap: Linguistic divergence and adaptation strategies in human-llm assistant vs. human-human interactions. arXiv preprint arXiv:2510.02645

  25. [33]

    Wenting Zhao, Xiang Ren, Jack Hessel, Claire Cardie, Yejin Choi, and Yuntian Deng. 2024. Wildchat: 1m chatgpt interaction logs in the wild. In The Twelfth International Conference on Learning Representations

  26. [34]

    online" 'onlinestring :=

    ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list...

  27. [35]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...

Pith tools

Reviewed June 29, 2026 · model on record in the stance chip above.