{"id":"c428a443-d792-410b-9b91-2e2417478af8","arxiv_id":"2507.18315","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"A Namer-Matcher experiment with 56 people found that a disfluent speech agent was rated as more competent after interaction than a fluent one, while disfluency's effect on perspective-taking language stayed uncertain.","lead":"In a naming game with a speech agent, people who heard the agent say 'uh' and 'um' ended up rating it as more competent than people who heard a fluent agent, because the fluent agent's rating dropped over the interaction. The result hints that small speech imperfections can change how users perceive and design for voice assistants.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The claim that the disfluent agent was perceived as 'more competent' is not directly supported by the reported analysis, which lacks a post-interaction between-group contrast.","rationale":"The reader's verdict is CONDITIONAL, and I agree with that level. However, the most load-bearing gap I see is not the (also plausible) absence of a believability manipulation check but the mismatch between the reported inferential tests and the headline interpretation. The key interaction is robust, but the paper's own central phrasing—'participants perceived the disfluent agent as more competent'—requires a post-interaction group contrast that is never reported. This is a concrete, checkable omission from the existing dataset, whereas the reverse-WoZ manipulation concern would require a new experiment to fully resolve. I therefore treat the missing post-hoc contrast as the single most load-bearing concern. If the contrast is available and significant, the paper's central claim stands and the CONDITIONAL verdict can be lifted; if not, the abstract and discussion should be softened. The paper otherwise gives credit for clear reporting of priors, convergence diagnostics, and honestly wide credibility intervals on the secondary egocentric-language analysis.","tokens_in":12089,"tokens_out":14177,"duration_ms":138082,"concrete_test":"Obtain the existing PMQ competence/dependability data (or ask the authors to run the test) and compute: (1) pre- and post-task means and SDs by condition; (2) the simple main effect of Speech Agent at post-interaction using the error term from the reported mixed ANOVA; (3) the same contrast with pre-task score as a covariate. If the post-interaction group contrast is not significant at p<.05, or is sensitive to baseline adjustment, revise the abstract and Discussion to say the fluent condition declined while the disfluent condition stayed stable, rather than that the disfluent agent was rated more competent.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The paper's headline PMQ claim is that, after interaction, participants perceived the disfluent agent as more competent and dependable than the fluent agent. What is actually reported in §5.2.1 is a significant Time × Speech Agent interaction (F(1,54)=13.77, p<.001) plus within-group pre/post contrasts: the fluent condition decreased (t(54)=5.68, p<.001) and the disfluent condition did not change (t(54)=0.22, p=.830). No simple main effect of Speech Agent at post-interaction is reported, and no pre-task means are given. A significant interaction can occur without a significant between-group difference at either time point, particularly if pre-task baselines are not perfectly balanced. The abstract's 'perceived the disfluent agent as more competent' and the Discussion's 'disfluent agents positively impact partner models' therefore go beyond the reported statistics; the supported statement is narrower: the fluent group's competence and dependability ratings declined over the task while the disfluent group's remained stable. This is an internal support gap rather than an external validity concern, and it is checkable from the existing data.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports an online Namer-Matcher experiment investigating whether a speech agent's disfluent speech (initial \"uh\"/\"um\" hesitation markers) affects users' partner models and perspective-taking. Fifty-six participants were analyzed across fluent (N=30) and disfluent (N=26) conditions. The authors report a significant Time by Speech Agent interaction on the PMQ Competence and Dependability subscale, driven by a decline in the fluent condition while the disfluent condition remained stable; no significant effects were found for Human-Likeness or Communicative Flexibility. Scalar modifier use was strongly affected by visual perspective, replicating prior work, but the effect of agent fluency on modifier use had wide credibility intervals and is interpreted cautiously. The Discussion interprets the PMQ result as showing that disfluent agents are perceived as more competent and that partner models are dynamic, while the perspective-taking results are framed as ambiguous evidence of egocentric versus audience-design strategies.","tokens_in":12329,"tokens_out":4907,"duration_ms":47120,"significance":"If the PMQ claim were fully supported by the reported statistics, the paper would make a useful contribution to the growing literature on partner models in human-machine dialogue, showing that a low-level speech design feature can update competence judgments over a short interaction. The paper also usefully replicates the perspective-condition effects of Peña et al. (2023) in an online setting and is honest about the wide credibility intervals for the fluency effect. The pre/post design with Bonferroni correction is a strength, as is the explicit discussion of ethical implications of using disfluency to influence perceived competence. However, the central PMQ conclusion currently goes beyond what the reported analyses demonstrate, and the absence of a manipulation check for the simulated real-time partner leaves an important interpretive gap.","major_comments":[{"comment":"The claim in the abstract and §6.1 that 'participants perceived the disfluent agent as more competent' and that 'disfluent agents positively impact partner models' is not supported by the statistics reported in §5.2.1. The reported results are a significant Time × Speech Agent interaction (F(1,54)=13.77, p<.001) with within-group pre/post contrasts showing a decline in the fluent condition (t(54)=5.68, p<.001) and no change in the disfluent condition (t(54)=0.22, p=.830); no post-interaction between-group simple effect and no pre-task means are reported. An interaction can be driven entirely by one group's change over time without a between-group difference at either time point. Please report the pre/post means with confidence intervals and a post-hoc contrast of Speech Agent at Post (or an equivalent simple-effects analysis), and revise the abstract, §5.2.1, and §6.1 to state exactly what the data support. The currently supported claim is that the fluent agent's ratings declined while the disfluent agent's ratings remained stable.","section":"§5.2.1, §6.1, Abstract"},{"comment":"The Communicative Flexibility subscale shows pre-task Cronbach's α = .53, below conventional reliability thresholds, yet §5.2.1 reports no statistically significant effects and §6.1 treats this as evidence of absence. The low reliability of the pre-task subscale substantially weakens any null conclusion for that subscale; report the analysis with appropriate caution, or omit strong null claims. Relatedly, the sample count is inconsistent: §3.1 reports 61 participants assigned to Fluent (N=30) and Disfluent (N=26) conditions, which sum to 56, while §5.2.1 and Fig. 2 say N=56. The paper should reconcile these numbers and state the effective N per analysis.","section":"§3.3.1 and §5.2.1"},{"comment":"The procedure relies on the reverse-Wizard-of-Oz framing with a 30-second 'Finding partner...' loading screen to 'support the illusion of a real-time partner,' but no manipulation check is reported to verify that participants believed they were interacting with an independent, real-time conversational partner. If participants did not perceive the interaction as live dialogue, the PMQ ratings and the perspective-taking measures could reflect reactions to pre-recorded speech rather than to a dialogue partner, which would weaken the central interpretation. Please add a manipulation check (e.g., post-task awareness items) or explicitly acknowledge in §6.3 that the believability of the simulated partner was unverified.","section":"§4"},{"comment":"The reporting of the fluency × perspective interaction appears internally inconsistent. §5.2.2 states that 'A strong difference was not observed between Common Ground and Privileged Ground for either the fluent (b = 0.68, SE = 0.35, 95% CrI [0.02, 1.39]) or the disfluent condition (b = 0.68, SE = 0.32, 95% CrI [0.06, 1.32])'; the identical point estimates for two different conditions suggest a copy-paste error, and the 95% credible intervals shown actually exclude zero, which is hard to reconcile with the text 'strong difference was not observed.' In addition, H3 predicted reduced scalar-modifier use in the privileged ground with the disfluent agent, but the reported tendency is in the opposite direction (increased use), and the Discussion reinterprets this as a possible audience-design strategy; the paper should explicitly state that H3 was not supported and provide the actual interaction estimates for the fluency × perspective terms (including the privileged-ground × disfluent contrast) rather than only the common-ground/privileged-ground simple effects.","section":"§5.2.2 and §2.3 (H3)"}],"minor_comments":[{"comment":"The final two sentences, 'Interaction with disfluent speech agents appears to increase egocentric communication in comparison to fluent agents. Although the wide credibility intervals mean this effect is not clear-cut,' form a sentence fragment and should be merged into a single hedged statement consistent with §5.2.2.","section":"Abstract"},{"comment":"The term 'reverse Wizard of Oz' usually implies a hidden human operator, whereas the method here uses pre-recorded TTS audio; please define the term or use a different label to avoid confusion.","section":"§3.2.3"},{"comment":"The scalar-modifier plot would be more informative with raw counts or error bars that correspond to the reported credibility intervals; the current mean/SE display does not convey the wide uncertainty emphasized in the text.","section":"Fig. 3"},{"comment":"'utternances' is a typo for 'utterances'; the Discussion also alternates between 'egocentric' and 'audience design' interpretations without a clear adjudication criterion, so a short decision rule or a proposed follow-up design would help the reader.","section":"§6.2"},{"comment":"The Bayesian priors for perspective effects are taken from Peña et al. (2023), a prior study by the same group; this is not circular because the fluency effect is not included in those priors, but a sensitivity analysis with weakly informative priors would strengthen the claim that the perspective effects are robust.","section":"§5.1"}],"recommendation":"major_revision","confidential_remarks":"The self-citation to Peña et al. (2023) is legitimate as prior work; the more pressing issue is the gap between the abstract's between-group claim and the reported interaction/within-group statistics. I do not see evidence of misconduct, but the manuscript needs a revision that brings the central claims in line with the analyses and resolves the sample-count inconsistency."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: the disfluency result is a real contribution to CUI/HMD, but the paper's headline sentence is stronger than its own statistics. The Time × Speech Agent interaction is solid, and the within-group contrasts show the fluent group's competence/dependability ratings dropped while the disfluent group stayed flat. That is not the same as showing the disfluent agent was rated significantly higher after interaction. No post-interaction simple main effect or pre-task means are reported, so the abstract's 'perceived the disfluent agent as more competent' goes beyond what is demonstrated. This is a narrow, fixable gap; the data can presumably answer it.\n\nWhat is genuinely new: combining a fluent vs disfluent TTS manipulation with the Partner Modelling Questionnaire and the Namer-Matcher perspective-taking task. No prior cited work did that. The study cleanly replicates the existing egocentric/allocentric scalar-modifier pattern in HMD, which is useful confirmation. The Bayesian handling of the modifier data is reasonable, with priors taken from Peña et al. (2023); that is same-group citation but not circular, since the fluency effect is not baked into the priors. The authors are also honest about the wide credibility intervals on the fluency effect, and the PMQ scales mostly show good reliability.\n\nSoft spots, in order. First, the missing post-hoc between-group contrast. Second, no manipulation check for the reverse-Wizard-of-Oz framing: the 30-second 'Finding partner…' loading screen is doing real work in the design, and without a check we cannot tell whether participants treated the agent as a live partner or as a canned voice. Third, the small sample (N=56 after exclusions) and the one-off online task limit how much weight to put on the null effects. Fourth, the communicative flexibility subscale had low pre-task internal consistency (α = .53), so null results on that subscale should be read cautiously. No data or code are provided, which makes the missing simple-effect test more annoying.\n\nWho this is for: CUI/HCI researchers and speech designers thinking about how small voice characteristics shape user perceptions and language. It deserves serious refereeing; the central design insight is useful even if the effect size and generalizability are uncertain. I would send it to review and ask for the follow-up analysis, a manipulation check, and ideally data availability.","headline":"The disfluent-agent competence finding is new and worth taking seriously, but the abstract overstates what the reported statistics show: the data support a stability-vs-decline pattern, not a directly tested post-interaction group difference.","tokens_in":12843,"tokens_out":2449,"would_cite":true,"duration_ms":25565,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that a speech agent's disfluency—hesitations like 'uh' and 'um'—changes users' partner models: after interaction, users rated the disfluent agent as more competent and dependable, while ratings of the fluent agent fell.","keywords":["speech disfluency","partner models","perspective taking","audience design","scalar modifiers","conversational agents","human-machine dialogue"],"falsifier":"Give participants the same pre-recorded fluent and disfluent speech but tell half of them the voice is pre-recorded and the other half that it is a live partner; if the competence-rating interaction disappears when the voice is known to be canned, the reverse-Wizard-of-Oz framing is load-bearing. Alternatively, include a post-task question asking whether the agent seemed to respond in real time and test whether non-believers show the same rating pattern.","tokens_in":11926,"feed_emoji":"🗣️","tokens_out":6113,"duration_ms":59357,"temperature":0.7,"pith_summary":"Speech disfluencies—the 'uh' and 'um' hesitations of natural speech—are known to carry information in human conversation, but their role in dialogue with machines is largely untested. This paper asks whether giving a speech agent disfluent utterances changes how users model the partner they are talking to, and whether it shifts the perspective users take when describing objects. Using a Namer-Matcher task with visual common ground and privileged ground, the authors show that competence and dependability ratings dropped after interaction for the fluent agent but stayed stable for the disfluent one, a significant Time by Speech Agent interaction. Their second finding, that disfluent agents may increase egocentric scalar-modifier use, is reported cautiously because credibility intervals are wide. A sympathetic reader would take the paper's contribution as evidence that small utterance-level design choices can reshape user perceptions and language production in human-machine dialogue.","feed_headline":"Disfluent speech agents keep user trust; fluent agents lose it","feed_subtitle":"Synthetic hesitations like 'uh' and 'um' keep an agent's competence rating steady while polished voices lose ground.","key_machinery":"The central machinery is the combination of (1) the Partner Modelling Questionnaire (PMQ), a self-report measure of perceived competence and dependability, human-likeness, and communicative flexibility; (2) the online Namer-Matcher task, which varies visual perspective so that a target object has no competitor, a competitor in common ground, or a competitor in privileged ground visible only to the user; and (3) pre-recorded speech constructed with initial 'uh' and 'um' hesitation markers, presented through a reverse Wizard-of-Oz framing to appear as a real-time agent. The PMQ supplies the outcome measure for partner models, the task supplies the behavioral measure of egocentric versus allocentric reference production, and the disfluency manipulation is the design variable whose effect the paper tests.","core_discovery":"The paper's central discovery is that disfluent speech changes partner models. Sixty-one participants rated the agent before and after a Namer-Matcher task; those who heard fluent descriptions downgraded the agent's competence and dependability afterward (t(54)=5.68, p<.001), while those who heard disfluent descriptions kept their ratings stable (t(54)=0.22, p=.830), producing a significant Time by Speech Agent interaction (F(1,54)=13.77, p<.001). The paper also replicates the finding that users use more scalar modifiers ('small') when a larger competitor is present in either common or privileged ground, and finds tentative evidence of more egocentric modifier use with the disfluent agent, while explicitly noting the wide credibility intervals make that effect uncertain.","pith_inferences":["A testable extension the paper leaves implicit: if disfluency signals task-state awareness, its competence benefit should grow as the task becomes more ambiguous; an experiment varying grid ambiguity could check that.","The fluent-agent drop may reflect an expectation-disconfirmation effect: users expect polished voices and recalibrate after a task in which the agent's knowledge of the visual world is imperfect. Measuring expectations separately from experience would separate these accounts.","The lack of a manipulation check on the 'Finding partner...' loading screen means a replication should ask whether participants believed they were talking to a real-time partner; the partner-model result may depend on that belief.","If the egocentric-modifier trend is real, disfluency could be acting like a licensing cue: users infer the agent can handle more informative descriptions, so they put less effort into perspective-taking. A follow-up with eye-tracking or referential error rates could test this mechanism."],"forward_implications":["If the PMQ result holds, designers who add naturalistic hesitations to synthetic voices may avoid a post-interaction drop in perceived competence and dependability that fluent voices appear to suffer.","The perspective-taking result reinforces that users in human-machine dialogue adapt their referring expressions to visual context, using modifiers to disambiguate competitors in both shared and privileged views.","The paper's reading of its own data implies that disfluency may signal task-state awareness to users, making an agent seem more attuned to ambiguity—a design cue rather than a defect.","Because the disfluency effect on egocentric modifier use was not conclusive, the paper supports the narrower claim that fluency of the agent's voice shapes partner models, while its effect on language production requires further evidence.","In trust-sensitive settings, the authors warn, using disfluency to raise perceived competence could raise ethical questions about transparency and user autonomy."],"supporting_citations":[{"why":"Supplies the online Namer-Matcher task, the three visual-perspective conditions, the Bayesian prior estimates, and the baseline audience-design and egocentrism findings being extended.","marker":"[39]"},{"why":"Provides the Partner Modelling Questionnaire, the pre/post measure of competence, dependability, human-likeness, and communicative flexibility.","marker":"[19]"},{"why":"Defines the partner-model construct—users' mental representations of a dialogue partner's communicative competence—that the study measures.","marker":"[18]"},{"why":"Establishes the communicative function of 'uh' and 'um' in spontaneous speech, the theoretical basis for predicting disfluency effects.","marker":"[14]"},{"why":"Supports the claim that users engage in audience design with computers and supplies the reverse Wizard-of-Oz simulation approach.","marker":"[8]"},{"why":"Shows that another speech design feature, accent, changes lexical choice in human-machine dialogue, motivating disfluency as a similar design variable.","marker":"[15]"}],"fun_headline_variants":["Stuttering saves face: disfluent agents dodge competence drop","Synthetic 'uh' keeps agent credibility, fluency backfires","Disfluent AI more competent? Users penalize smooth talkers","Hesitant bots win competence ratings in partner tests","Um, actually: disfluent agents retain user trust"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper assumes participants accepted the pre-recorded voice as a genuine, real-time conversational partner; the 30-second loading screen was crafted for that illusion, but no manipulation check is reported to confirm it worked.","fun_headline_variants_meta":{"raw":{"variants":["Stuttering saves face: disfluent agents dodge competence drop","Synthetic 'uh' keeps agent credibility, fluency backfires","Disfluent AI more competent? Users penalize smooth talkers","Hesitant bots win competence ratings in partner tests","Um, actually: disfluent agents retain user trust"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000545,"raw_usage":{"total_tokens":2583,"prompt_tokens":896,"completion_tokens":1687,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":512,"completion_tokens_details":{"reasoning_tokens":1602}},"tokens_in":512,"tokens_out":1687,"duration_ms":11917,"temperature":1.0,"reasoning_tokens":1602,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T18:15:01.620956+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Give participants the same pre-recorded fluent and disfluent speech but tell half of them the voice is pre-recorded and the other half that it is a live partner; if the competence-rating interaction disappears when the voice is known to be canned, the reverse-Wizard-of-Oz framing is load-bearing. Alternatively, include a post-task question asking whether the agent seemed to respond in real time and test whether non-believers show the same rating pattern.","supporting_citations":[{"cited_title":"The Partner Modelling Questionnaire: A validated self-report measure of perceptions toward machines as dialogue partners","cited_arxiv_id":"2308.07164","evidence_quote":"Provides the Partner Modelling Questionnaire, the pre/post measure of competence, dependability, human-likeness, and communicative flexibility."},{"cited_title":"Branigan, Martin J","cited_arxiv_id":null,"evidence_quote":"Supports the claim that users engage in audience design with computers and supplies the reverse Wizard-of-Oz simulation approach."}],"review_version":2}