Pith. sign in

REVIEW 4 major objections 5 minor 47 references

Talking to...uh...um...Machines: The Impact of Disfluent Speech Agents on Partner Models and Perspective Taking

T0 review · 4 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read This paper claims that a speech agent's disfluency—hesitations like 'uh' and 'um'—changes users' partner models: after interaction, users rated the disfluent agent as more competent and dependable, while ratings of the fluent agent fell.

desk verdict The disfluent-agent competence finding is new and worth taking seriously, but the abstract overstates what the reported statistics show: the data support a stability-vs-decline pattern, not a directly tested post-interaction group difference. read the letter →

arxiv 2507.18315 v1 pith:KTASWYLC submitted 2025-07-24 cs.HC

classification cs.HC
keywords speechdisfluencypartnermodelsperspectivetakingaudiencedesignscalarmodifiersconversationalagentshuman-machinedialogue
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Speech disfluencies—the 'uh' and 'um' hesitations of natural speech—are known to carry information in human conversation, but their role in dialogue with machines is largely untested. This paper asks whether giving a speech agent disfluent utterances changes how users model the partner they are talking to, and whether it shifts the perspective users take when describing objects. Using a Namer-Matcher task with visual common ground and privileged ground, the authors show that competence and dependability ratings dropped after interaction for the fluent agent but stayed stable for the disfluent one, a significant Time by Speech Agent interaction. Their second finding, that disfluent agents may increase egocentric scalar-modifier use, is reported cautiously because credibility intervals are wide. A sympathetic reader would take the paper's contribution as evidence that small utterance-level design choices can reshape user perceptions and language production in human-machine dialogue.

What carries the argument

The central machinery is the combination of (1) the Partner Modelling Questionnaire (PMQ), a self-report measure of perceived competence and dependability, human-likeness, and communicative flexibility; (2) the online Namer-Matcher task, which varies visual perspective so that a target object has no competitor, a competitor in common ground, or a competitor in privileged ground visible only to the user; and (3) pre-recorded speech constructed with initial 'uh' and 'um' hesitation markers, presented through a reverse Wizard-of-Oz framing to appear as a real-time agent. The PMQ supplies the outcome measure for partner models, the task supplies the behavioral measure of egocentric versus allocentric reference production, and the disfluency manipulation is the design variable whose effect the paper tests.

What would settle it

Give participants the same pre-recorded fluent and disfluent speech but tell half of them the voice is pre-recorded and the other half that it is a live partner; if the competence-rating interaction disappears when the voice is known to be canned, the reverse-Wizard-of-Oz framing is load-bearing. Alternatively, include a post-task question asking whether the agent seemed to respond in real time and test whether non-believers show the same rating pattern.

Watch

Extended reading notes

Core claim

The paper's central discovery is that disfluent speech changes partner models. Sixty-one participants rated the agent before and after a Namer-Matcher task; those who heard fluent descriptions downgraded the agent's competence and dependability afterward (t(54)=5.68, p<.001), while those who heard disfluent descriptions kept their ratings stable (t(54)=0.22, p=.830), producing a significant Time by Speech Agent interaction (F(1,54)=13.77, p<.001). The paper also replicates the finding that users use more scalar modifiers ('small') when a larger competitor is present in either common or privileged ground, and finds tentative evidence of more egocentric modifier use with the disfluent agent, while explicitly noting the wide credibility intervals make that effect uncertain.

Load-bearing premise

The paper assumes participants accepted the pre-recorded voice as a genuine, real-time conversational partner; the 30-second loading screen was crafted for that illusion, but no manipulation check is reported to confirm it worked.

Editorial extensions

If this is right

  • If the PMQ result holds, designers who add naturalistic hesitations to synthetic voices may avoid a post-interaction drop in perceived competence and dependability that fluent voices appear to suffer.
  • The perspective-taking result reinforces that users in human-machine dialogue adapt their referring expressions to visual context, using modifiers to disambiguate competitors in both shared and privileged views.
  • The paper's reading of its own data implies that disfluency may signal task-state awareness to users, making an agent seem more attuned to ambiguity—a design cue rather than a defect.
  • Because the disfluency effect on egocentric modifier use was not conclusive, the paper supports the narrower claim that fluency of the agent's voice shapes partner models, while its effect on language production requires further evidence.
  • In trust-sensitive settings, the authors warn, using disfluency to raise perceived competence could raise ethical questions about transparency and user autonomy.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A testable extension the paper leaves implicit: if disfluency signals task-state awareness, its competence benefit should grow as the task becomes more ambiguous; an experiment varying grid ambiguity could check that.
  • The fluent-agent drop may reflect an expectation-disconfirmation effect: users expect polished voices and recalibrate after a task in which the agent's knowledge of the visual world is imperfect. Measuring expectations separately from experience would separate these accounts.
  • The lack of a manipulation check on the 'Finding partner...' loading screen means a replication should ask whether participants believed they were talking to a real-time partner; the partner-model result may depend on that belief.
  • If the egocentric-modifier trend is real, disfluency could be acting like a licensing cue: users infer the agent can handle more informative descriptions, so they put less effort into perspective-taking. A follow-up with eye-tracking or referential error rates could test this mechanism.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper reports an online Namer-Matcher experiment investigating whether a speech agent's disfluent speech (initial "uh"/"um" hesitation markers) affects users' partner models and perspective-taking. Fifty-six participants were analyzed across fluent (N=30) and disfluent (N=26) conditions. The authors report a significant Time by Speech Agent interaction on the PMQ Competence and Dependability subscale, driven by a decline in the fluent condition while the disfluent condition remained stable; no significant effects were found for Human-Likeness or Communicative Flexibility. Scalar modifier use was strongly affected by visual perspective, replicating prior work, but the effect of agent fluency on modifier use had wide credibility intervals and is interpreted cautiously. The Discussion interprets the PMQ result as showing that disfluent agents are perceived as more competent and that partner models are dynamic, while the perspective-taking results are framed as ambiguous evidence of egocentric versus audience-design strategies.

Significance. If the PMQ claim were fully supported by the reported statistics, the paper would make a useful contribution to the growing literature on partner models in human-machine dialogue, showing that a low-level speech design feature can update competence judgments over a short interaction. The paper also usefully replicates the perspective-condition effects of Peña et al. (2023) in an online setting and is honest about the wide credibility intervals for the fluency effect. The pre/post design with Bonferroni correction is a strength, as is the explicit discussion of ethical implications of using disfluency to influence perceived competence. However, the central PMQ conclusion currently goes beyond what the reported analyses demonstrate, and the absence of a manipulation check for the simulated real-time partner leaves an important interpretive gap.

major comments (4)
  1. [§5.2.1, §6.1, Abstract] The claim in the abstract and §6.1 that 'participants perceived the disfluent agent as more competent' and that 'disfluent agents positively impact partner models' is not supported by the statistics reported in §5.2.1. The reported results are a significant Time × Speech Agent interaction (F(1,54)=13.77, p<.001) with within-group pre/post contrasts showing a decline in the fluent condition (t(54)=5.68, p<.001) and no change in the disfluent condition (t(54)=0.22, p=.830); no post-interaction between-group simple effect and no pre-task means are reported. An interaction can be driven entirely by one group's change over time without a between-group difference at either time point. Please report the pre/post means with confidence intervals and a post-hoc contrast of Speech Agent at Post (or an equivalent simple-effects analysis), and revise the abstract, §5.2.1, and §6.1 to state exactly what the data support. The currently supported claim is that the fluent agent's ratings declined while the disfluent agent's ratings remained stable.
  2. [§3.3.1 and §5.2.1] The Communicative Flexibility subscale shows pre-task Cronbach's α = .53, below conventional reliability thresholds, yet §5.2.1 reports no statistically significant effects and §6.1 treats this as evidence of absence. The low reliability of the pre-task subscale substantially weakens any null conclusion for that subscale; report the analysis with appropriate caution, or omit strong null claims. Relatedly, the sample count is inconsistent: §3.1 reports 61 participants assigned to Fluent (N=30) and Disfluent (N=26) conditions, which sum to 56, while §5.2.1 and Fig. 2 say N=56. The paper should reconcile these numbers and state the effective N per analysis.
  3. [§4] The procedure relies on the reverse-Wizard-of-Oz framing with a 30-second 'Finding partner...' loading screen to 'support the illusion of a real-time partner,' but no manipulation check is reported to verify that participants believed they were interacting with an independent, real-time conversational partner. If participants did not perceive the interaction as live dialogue, the PMQ ratings and the perspective-taking measures could reflect reactions to pre-recorded speech rather than to a dialogue partner, which would weaken the central interpretation. Please add a manipulation check (e.g., post-task awareness items) or explicitly acknowledge in §6.3 that the believability of the simulated partner was unverified.
  4. [§5.2.2 and §2.3 (H3)] The reporting of the fluency × perspective interaction appears internally inconsistent. §5.2.2 states that 'A strong difference was not observed between Common Ground and Privileged Ground for either the fluent (b = 0.68, SE = 0.35, 95% CrI [0.02, 1.39]) or the disfluent condition (b = 0.68, SE = 0.32, 95% CrI [0.06, 1.32])'; the identical point estimates for two different conditions suggest a copy-paste error, and the 95% credible intervals shown actually exclude zero, which is hard to reconcile with the text 'strong difference was not observed.' In addition, H3 predicted reduced scalar-modifier use in the privileged ground with the disfluent agent, but the reported tendency is in the opposite direction (increased use), and the Discussion reinterprets this as a possible audience-design strategy; the paper should explicitly state that H3 was not supported and provide the actual interaction estimates for the fluency × perspective terms (including the privileged-ground × disfluent contrast) rather than only the common-ground/privileged-ground simple effects.
minor comments (5)
  1. [Abstract] The final two sentences, 'Interaction with disfluent speech agents appears to increase egocentric communication in comparison to fluent agents. Although the wide credibility intervals mean this effect is not clear-cut,' form a sentence fragment and should be merged into a single hedged statement consistent with §5.2.2.
  2. [§3.2.3] The term 'reverse Wizard of Oz' usually implies a hidden human operator, whereas the method here uses pre-recorded TTS audio; please define the term or use a different label to avoid confusion.
  3. [Fig. 3] The scalar-modifier plot would be more informative with raw counts or error bars that correspond to the reported credibility intervals; the current mean/SE display does not convey the wide uncertainty emphasized in the text.
  4. [§6.2] 'utternances' is a typo for 'utterances'; the Discussion also alternates between 'egocentric' and 'audience design' interpretations without a clear adjudication criterion, so a short decision rule or a proposed follow-up design would help the reader.
  5. [§5.1] The Bayesian priors for perspective effects are taken from Peña et al. (2023), a prior study by the same group; this is not circular because the fluency effect is not included in those priors, but a sensitivity analysis with weakly informative priors would strengthen the claim that the perspective effects are robust.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity; the Bayesian priors are transparent external evidence and the fluency effect is not encoded in them.

full rationale

The paper is an empirical study rather than a derivation chain, and its central claims are estimated from newly collected participant data. The only author-supplied prior information appears in Section 5.1, where the Bayesian priors for scalar modifier use are taken from Peña et al. (2023). Those priors concern the Perspective-condition effects (privileged and common ground relative to one target), not the manipulation of interest, agent fluency. The fluency main effect and its interactions with Perspective are not assigned informative priors from the same group's earlier work, so the paper's headline fluency-related results are not forced by prior input. The cited prior study is an externally falsifiable, previously published experiment using the same Namer-Matcher task, so importing its estimates is legitimate Bayesian evidence rather than a self-referential input. The Partner Modelling Questionnaire analysis is a standard mixed ANOVA on pre/post ratings with no fitted parameter later renamed as a prediction; the competence and dependability interaction is derived from measured ratings, not from an equation that encodes the conclusion. The paper also reports a disfluency effect on scalar modifiers that is uncertain and contrary to the direction of its own H3 prediction, which further indicates the analysis was not constructed to guarantee the reported outcome. No quoted equation or fitted value makes any reported result true by construction, and no circular step can be exhibited under the required standard. The paper is therefore self-contained against external benchmarks, and the honest finding is no significant circularity.

Assumptions & free parameters 1 free parameters · 5 assumptions · 0 invented entities

The central PMQ finding requires no free parameters; the only imported numeric choices are the informative Bayesian priors for the scalar-modifier model, taken from the same group's earlier study. The behavioral interpretation rests on standard psycholinguistic assumptions about common ground and scalar modifier use.

free parameters (1)
  • Informative normal prior means for perspective effects (from Peña et al. 2023) = PG: B=7.88, SE=1.15; CG: B=9.96, SE=1.19
    Chosen from the prior study rather than estimated in this paper; used as priors in the Bayesian GLMM for scalar modifier use. These values do not determine the fluency effect, which is the paper's uncertain secondary claim.
assumptions (5)
  • domain assumption The Partner Modelling Questionnaire (Doyle et al. 2025) validly measures partner models of machine dialogue partners.
    The PMQ subscales for competence, human-likeness, and flexibility are interpreted as partner model dimensions; reliabilities are reported except the pre-task communicative flexibility subscale, which is low (alpha=.53).
  • domain assumption The Namer-Matcher task (Peña et al. 2023) validly creates common ground, privileged ground, and one-target perspective conditions.
    Visual perspective is manipulated through white versus grey backgrounds, following the cited prior task; the authors assume this manipulation produces the intended knowledge asymmetries.
  • domain assumption Scalar modifier use ('small' etc.) is a valid behavioral index of perspective-taking and egocentric or allocentric language production.
    The coding scheme treats modifier use as evidence of speakers adjusting to competitor presence; this interpretation is carried over from prior psycholinguistic work.
  • domain assumption The reverse Wizard of Oz framing made the pre-recorded agent a believable, real-time dialogue partner.
    Section 4 describes the 30-second 'Finding partner...' loading screen to support the illusion, but no manipulation check is reported.
  • domain assumption The informative priors from Peña et al. (2023) are appropriate for the current sample and task.
    The priors come from the same Namer-Matcher paradigm with a different sample; the authors assume transferability, which the wide credibility intervals suggest is uncertain.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Talking to...uh...um...Machines: The Impact of Disfluent Speech Agents on Partner Models and Perspective Taking." pith.science (2026). https://pith.science/paper/KTASWYLC

@misc{pith2026250718315,
  author       = {Pith},
  title        = {Pith review of: Talking to...uh...um...Machines: The Impact of Disfluent Speech Agents on Partner Models and Perspective Taking},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/KTASWYLC}},
  note         = {Machine review of arXiv:2507.18315}
}
read the original abstract

Speech disfluencies play a role in perspective-taking and audience design in human-human communication (HHC), but little is known about their impact in human-machine dialogue (HMD). In an online Namer-Matcher task, sixty-one participants interacted with a speech agent using either fluent or disfluent speech. Participants completed a partner-modelling questionnaire (PMQ) both before and after the task. Post-interaction evaluations indicated that participants perceived the disfluent agent as more competent, despite no significant differences in pre-task ratings. However, no notable differences were observed in assessments of conversational flexibility or human-likeness. Our findings also reveal evidence of egocentric and allocentric language production when participants interact with speech agents. Interaction with disfluent speech agents appears to increase egocentric communication in comparison to fluent agents. Although the wide credibility intervals mean this effect is not clear-cut. We discuss potential interpretations of this finding, focusing on how disfluencies may impact partner models and language production in HMD.

Figures

Figures reproduced from arXiv: 2507.18315 by the authors.

Figure 1
Figure 1. Image replicated from Peña et al. (2023) showing experiment perspective conditions. [PITH_FULL_IMAGE:figures/full_fig_p005_1.png] view at source ↗
Figure 2
Figure 2. Bar plots portraying changes in PMQ score for each of the three subscales. (a) Competence and Dependability, (b) Human [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗
Figure 3
Figure 3. Mean and standard error of proportion of scalar modifiers used between partner conditions across perspective conditions, N= [PITH_FULL_IMAGE:figures/full_fig_p008_3.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

47 extracted references · 23 canonical work pages

  1. [1]

    Sungeun An, Robert Moore, Eric Young Liu, and Guang-Jie Ren. 2021. Recipient Design for Conversational Agents: Tailoring Agent’s Utterance to User’s Knowledge. In Proceedings of the 3rd Conference on Conversational User Interfaces (CUI ’21) . Association for Computing Machinery, New York, NY, USA, 1–5. doi:10.1145/3469595.3469625

  2. [2]

    Arnold, Carla L

    Jennifer E. Arnold, Carla L. Hudson Kam, and Michael K. Tanenhaus. 2007. If you say thee uh you are describing something hard: The on-line attribution of disfluency during reference comprehension. Journal of Experimental Psychology: Learning, Memory, and Cognition 33, 5 (2007), 914–930. doi:10.1037/0278-7393.33.5.914 Manuscript submitted to ACM Talking to...

  3. [3]

    Arnold, Michael K

    Jennifer E. Arnold, Michael K. Tanenhaus, Rebecca J. Altmann, and Maria Fagnano. 2004. The Old and Thee, uh, New: Disfluency and Reference Resolution. Psychological Science 15, 9 (Sept. 2004), 578–582. doi:10.1111/j.0956-7976.2004.00723.x Publisher: SAGE Publications Inc

  4. [4]

    Barr and Mandana Seyfeddinipur

    Dale J. Barr and Mandana Seyfeddinipur. 2010. The role of fillers in listener attributions for speaker disfluency. Language and Cognitive Processes 25, 4 (May 2010), 441–455. doi:10.1080/01690960903047122 Publisher: Routledge _eprint: https://doi.org/10.1080/01690960903047122

  5. [5]

    Allan Bell. 1984. Language style as audience design. Language in Society 13, 2 (June 1984), 145–204. doi:10.1017/S004740450001037X

  6. [6]

    Simon Betz, Birte Carlmeyer, Petra Wagner, and Britta Wrede. 2018. Interactive Hesitation Synthesis: Modelling and Evaluation. Multimodal Technologies and Interaction 2, 1 (March 2018), 9. doi:10.3390/mti2010009 Number: 1 Publisher: Multidisciplinary Digital Publishing Institute

  7. [7]

    Dan Bohus and Eric Horvitz. 2014. Managing Human-Robot Engagement with Forecasts and... um... Hesitations. InProceedings of the 16th International Conference on Multimodal Interaction (ICMI ’14) . Association for Computing Machinery, New York, NY, USA, 2–9. doi:10.1145/2663204.2663241

  8. [8]

    Branigan, Martin J

    Holly P. Branigan, Martin J. Pickering, Jamie Pearson, Janet F. McLean, and Ash Brown. 2011. The role of beliefs in lexical alignment: Evidence from dialogs with humans and computers. Cognition 121, 1 (Oct. 2011), 41–57. doi:10.1016/j.cognition.2011.05.011

Show all 47 references
  1. [9]

    Brennan and Joy E

    Susan E. Brennan and Joy E. Hanna. 2009. Partner-Specific Adaptation in Dialog. Topics in Cognitive Science 1, 2 (2009), 274–291. doi:10.1111/j.1756- 8765.2009.01019.x _eprint: https://onlinelibrary.wiley.com/doi/pdf/10.1111/j.1756-8765.2009.01019.x

  2. [10]

    S. E. Brennan and M. Williams. 1995. The Feeling of Another’s Knowing: Prosody and Filled Pauses as Cues to Listeners about the Metacognitive States of Speakers. Journal of Memory and Language 34, 3 (June 1995), 383–398. doi:10.1006/jmla.1995.1017

  3. [11]

    Xinyi Chen, Andreas Liesenfeld, Shiyue Li, and Yao Yao. 2022. Effects of Filled Pauses on Memory Recall in Human-Robot Interaction in Mandarin Chinese. In Engineering Psychology and Cognitive Ergonomics , Don Harris and Wen-Chin Li (Eds.). Springer International Publishing, Ch...

  4. [12]

    Herbert H. Clark. 1992. Arenas of Language Use. University of Chicago Press

  5. [13]

    Herbert H. Clark. 1996. Using Language. Cambridge University Press, Cambridge. doi:10.1017/CBO9780511620539

  6. [14]

    Clark and Jean E

    Herbert H. Clark and Jean E. Fox Tree. 2002. Using uh and um in spontaneous speaking. Cognition 84, 1 (May 2002), 73–111. doi:10.1016/S0010- 0277(02)00017-3

  7. [15]

    Cowan, Philip Doyle, Justin Edwards, Diego Garaialde, Ali Hayes-Brady, Holly P

    Benjamin R. Cowan, Philip Doyle, Justin Edwards, Diego Garaialde, Ali Hayes-Brady, Holly P. Branigan, João Cabral, and Leigh Clark. 2019. What’s in an accent? the impact of accented synthetic speech on lexical choice in human-machine dialogue. InProceedings of the 1st Internat...

  8. [16]

    What can i help you with?

    Benjamin R. Cowan, Nadia Pantidi, David Coyle, Kellie Morrissey, Peter Clarke, Sara Al-Shehri, David Earley, and Natasha Bandeira. 2017. "What can i help you with?": infrequent users’ experiences of intelligent personal assistants. In Proceedings of the 19th International Conf...

  9. [17]

    Judit Dombi, Tetyana Sydorenko, and Veronika Timpe-Laughlin. 2022. Common ground, cooperation, and recipient design in human-computer interactions. Journal of Pragmatics 193 (May 2022), 4–20. doi:10.1016/j.pragma.2022.03.001

  10. [18]

    Philip R Doyle, Leigh Clark, and Benjamin R. Cowan. 2021. What Do We See in Them? Identifying Dimensions of Partner Models for Speech Interfaces Using a Psycholexical Approach. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems (CHI ’21) . Associat...

  11. [19]

    Doyle, Iona Gessinger, Justin Edwards, Leigh Clark, Odile Dumbleton, Diego Garaialde, Daniel Rough, Anna Bleakley, Holly P

    Philip R. Doyle, Iona Gessinger, Justin Edwards, Leigh Clark, Odile Dumbleton, Diego Garaialde, Daniel Rough, Anna Bleakley, Holly P. Branigan, and Benjamin R. Cowan. 2025. The Partner Modelling Questionnaire: A validated self-report measure of perceptions toward machines as d...

  12. [20]

    Ylva Ferstl and Rachel McDonnell. 2018. Investigating the use of recurrent motion modelling for speech gesture generation. In Proceedings of the 18th International Conference on Intelligent Virtual Agents . ACM, Sydney NSW Australia, 93–98. doi:10.1145/3267851.3267898

  13. [21]

    Flavell, Barbara A

    John H. Flavell, Barbara A. Everett, Karen Croft, and Eleanor R. Flavell. 1981. Young children’s knowledge about visual perception: Further evidence for the Level 1–Level 2 distinction. Developmental Psychology 17, 1 (1981), 99–103. doi:10.1037/0012-1649.17.1.99 Place: US Publ...

  14. [22]

    Alexia Galati and Susan E. Brennan. 2010. Attenuating information in spoken communication: For the speaker, or for the addressee? Journal of Memory and Language 62, 1 (Jan. 2010), 35–51. doi:10.1016/j.jml.2009.09.002

  15. [23]

    Iona Gessinger, Katie Seaborn, Madeleine Steeds, and Benjamin R. Cowan. 2025. ChatGPT and me: First-time and experienced users’ perceptions of ChatGPT’s communicative ability as a dialogue partner. International Journal of Human-Computer Studies 194 (Feb. 2025), 103400. doi:10...

  16. [24]

    Ella Glikson and Anita Woolley. 2020. Human trust in artificial intelligence: Review of empirical research. Academy of Management Annals (in press). The Academy of Management Annals (April 2020)

  17. [25]

    Hawkins, Hyowon Gweon, and Noah D

    Robert D. Hawkins, Hyowon Gweon, and Noah D. Goodman. 2021. The Division of Labor in Communication: Speakers Help Lis- teners Account for Asymmetries in Visual Perspective. Cognitive Science 45, 3 (2021), e12926. doi:10.1111/cogs.12926 _eprint: https://onlinelibrary.wiley.com/...

  18. [26]

    Aike Horstmann, Clara Strathmann, Lea Lambrich, and Nicole Krämer. 2024. Communication Style Adaptation in Human-Computer Interaction: An Empirical Study on the Effects of a Voice Assistant’s Politeness and Machine-Likeness on People’s Communication Behavior During and After t...

  19. [27]

    Horton and Boaz Keysar

    William S. Horton and Boaz Keysar. 1996. When do speakers take into account common ground?Cognition 59, 1 (April 1996), 91–117. doi:10.1016/0010- 0277(96)81418-1

  20. [28]

    Cowan, and Donald Mcmillan

    Razan Jaber, Sabrina Zhong, Sanna Kuoppamäki, Aida Hosseini, Iona Gessinger, Duncan P Brumby, Benjamin R. Cowan, and Donald Mcmillan

  21. [29]

    Boaz Keysar. 2007. Communication and miscommunication: The role of egocentric processes. 4, 1 (March 2007), 71–84. doi:10.1515/IP.2007.004 Publisher: De Gruyter Mouton Section: Intercultural Pragmatics

  22. [30]

    Barr, Jennifer A

    Boaz Keysar, Dale J. Barr, Jennifer A. Balin, and Jason S. Brauner. 2000. Taking Perspective in Conversation: The Role of Mutual Knowledge in Comprehension. https://journals.sagepub.com/doi/abs/10.1111/1467-9280.00211

  23. [31]

    Katerina Koleva, Maurizio Vergari, Tanja Kojić, Sebastian Möller, and Jan-Niklas Voigt-Antons. 2024. Influence of Personality and Communication Behavior of a Conversational Agent on User Experience and Social Presence in Augmented Reality. doi:10.48550/arXiv.2403.09883 arXiv:2...

  24. [32]

    Tanya Kraljic and Susan E. Brennan. 2005. Prosodic disambiguation of syntactic structure: For the speaker or for the addressee?Cognitive Psychology 50, 2 (March 2005), 194–231. doi:10.1016/j.cogpsych.2004.08.002

  25. [33]

    Loy and Vera Demberg

    Jia E. Loy and Vera Demberg. 2023. Perspective Taking Reflects Beliefs About Partner Sophistication: Modern Computer Part- ners Versus Basic Computer and Human Partners. Cognitive Science 47, 12 (2023), e13385. doi:10.1111/cogs.13385 _eprint: https://onlinelibrary.wiley.com/do...

  26. [34]

    Like Having a Really Bad PA

    Ewa Luger and Abigail Sellen. 2016. "Like Having a Really Bad PA": The Gulf between User Expectation and Experience of Conversational Agents. In Proceedings of the 2016 CHI Conference on Human Factors in Computing Systems (CHI ’16) . Association for Computing Machinery, New Yo...

  27. [35]

    J¯ura Miniota, Siyang Wang, Jonas Beskow, Joakim Gustafson, Éva Székely, and André Pereiral. 2023. Hi robot, it’s not what you say, it’s how you say it. In 2023 32nd IEEE International Conference on Robot and Human Interactive Communication (RO-MAN) . 307–314. doi:10.1109/RO- ...

  28. [36]

    Clifford Nass, Youngme Moon, and Paul Carney. 1999. Are People Polite to Computers? Responses to Computer-Based Inter- viewing Systems. Journal of Applied Social Psychology 29, 5 (1999), 1093–1109. doi:10.1111/j.1559-1816.1999.tb00142.x _eprint: https://onlinelibrary.wiley.com...

  29. [37]

    Paola R Peña, Philip R Doyle, Diego Garaialde, Yunhan Wu, Rachel McDonnell, and Benjamin R Cowan. 2023. Human Speakers Help Machine Listeners To account For Visual Asymmetries in Dialogue. In Proceedings of the 5th International Conference on Conversational User Interfaces . 1–13

  30. [38]

    Paola R. Peña. 2025. Perceptions of Machine Partners Competence Influence Perspective-Taking in Speech Agent Communication. (2025). Manuscript in preparation

  31. [39]

    Peña, Philip Doyle, Justin Edwards, Diego Garaialde, Daniel Rough, Anna Bleakley, Leigh Clark, Anita Tobar Henriquez, Holly Branigan, Iona Gessinger, and Benjamin R

    Paola R. Peña, Philip Doyle, Justin Edwards, Diego Garaialde, Daniel Rough, Anna Bleakley, Leigh Clark, Anita Tobar Henriquez, Holly Branigan, Iona Gessinger, and Benjamin R. Cowan. 2023. Audience design and egocentrism in reference production during human-computer dialogue. I...

  32. [40]

    Pickering and Simon Garrod

    Martin J. Pickering and Simon Garrod. 2004. Toward a mechanistic psychology of dialogue. Behavioral and Brain Sciences 27, 2 (April 2004), 169–190. doi:10.1017/S0140525X04000056

  33. [41]

    Hartsuiker

    Aurélie Pistono, , and Robert J. Hartsuiker. 2021. Eye-movements can help disentangle mechanisms underlying disfluency. Lan- guage, Cognition and Neuroscience 36, 8 (Oct. 2021), 1038–1055. doi:10.1080/23273798.2021.1905166 Publisher: Routledge _eprint: https://doi.org/10.1080/...

  34. [42]

    Fischer, Stuart Reeves, and Sarah Sharples

    Martin Porcheron, Joel E. Fischer, Stuart Reeves, and Sarah Sharples. 2018. Voice Interfaces in Everyday Life. InProceedings of the 2018 CHI Conference on Human Factors in Computing Systems . ACM, Montreal QC Canada, 1–12. doi:10.1145/3173574.3174214

  35. [43]

    Minjin Rheu, Shin , Ji Youn, Peng , Wei, , and Jina Huh-Yoo. 2021. Systematic Review: Trust-Building Factors and Implications for Conversational Agent Design. International Journal of Human–Computer Interaction 37, 1 (Jan. 2021), 81–96. doi:10.1080/10447318.2020.1807710 Publis...

  36. [44]

    Rothwell, Valerie L

    Clayton D. Rothwell, Valerie L. Shalin, and Griffin D. Romigh. 2021. Comparison of Common Ground Models for Human–Computer Dialogue: Evidence for Audience Design. ACM Trans. Comput.-Hum. Interact. 28, 2 (April 2021), 9:1–9:35. doi:10.1145/3410876

  37. [45]

    Michael F. Schober. 1993. Spatial perspective-taking in conversation. Cognition 47, 1 (Jan. 1993), 1–24. doi:10.1016/0010-0277(93)90060-9

  38. [46]

    Hartsuiker, and Aurélie Pistono

    Kasper Van Craeyenest, Bram De Keersmaecker, Robert J. Hartsuiker, and Aurélie Pistono. 2025. Do filled pauses serve a . . . um. . . communicative function? A comparison of self-directed and social speech. Discourse Processes 0, 0 (2025), 1–14. doi:10.1080/0163853X.2025.246863...

  39. [2024]

    In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems (CHI ’24)

    Cooking With Agents: Designing Context-aware Voice Interaction. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems (CHI ’24). Association for Computing Machinery, New York, NY, USA, 1–13. doi:10.1145/3613904.3642183

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.