{"id":"3cd6a903-e2dd-4b7d-8214-0742f6cdaa44","arxiv_id":"2607.22939","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Speakers locally entrain to human partners but not to a digital agent in multiparty conversations; global entrainment is limited and cohort-dependent.","lead":"In group conversations with a digital agent, adults matched each other's pitch, emotion, and speaking style within each turn, but did not match the agent. Children adapted to the agent over time, while global entrainment among adults and families stayed weak.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Local entrainment result may be an artifact of shared agent prompts; Section 3.1 model has no prompt/topic control.","rationale":"The paper's headline claim is that adults locally entrain with other humans while not with the agent. The statistical evidence is the negative γ in Tables 2 and 3. The experimental design embeds a structural common-cause confound: within-turn pairs share the agent's prompt. This is not a minor omission—it threatens the causal attribution at the heart of the paper. The authors define a turn as an agent prompt plus responses; no analysis step removes prompt influence. Effect sizes align with the alternative explanation: semantic embeddings (Whisper) show strong effects, and handcrafted features like pitch and amplitude could also be affected if questions elicit similar affective responses. The absence of P2A effects does not help because comparing a human voice to a synthetic agent voice is dominated by voice identity, masking topic similarity. The proposed test—matching same prompts across sessions, or using prompt-distance covariates—would settle whether γ reflects genuine inter-speaker adaptation or merely within-turn topical/emotional coherence. If the effect fails, the central claim must be weakened; if it passes, the entrainment result is real. This concern aligns with the reader's weakest_assumption. It justifies keeping the verdict at CONDITIONAL: the dataset is a real contribution, but the primary analysis needs this control before the claim can be accepted. I therefore recommend no change from the reader's conditional verdict.","tokens_in":8071,"tokens_out":14519,"duration_ms":128338,"concrete_test":"Identify all turns in different sessions that share the same agent prompt (same dialog-tree node). Restrict the Section 3.1 regression to utterance pairs of type (a) same-turn/same-session (same prompt, live interaction) and (b) different-session/same-prompt (same prompt, no interaction). Re-estimate γ on this matched set. If the same-turn coefficient is no longer significantly negative (Bonferroni-corrected p ≥ 0.05) or drops by more than half relative to Tables 2–3, the reported local entrainment is explained by the common prompt. This requires that prompt identity be recorded; if it is not, the paper must add it or the finding remains uninterpretable.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central positive result—adult local entrainment in Section 3.2—is established by a regression (Section 3.1) that compares feature distances for utterance pairs inside the same turn vs across turns. By the paper's turn definition, a turn is 'one agent prompt followed by the participants' responses.' Thus two different human utterances in the same turn share that prompt verbatim, while cross-turn pairs have different prompts. The shared prompt is a common cause: it can make responses more similar in topic, emotional connotation, and even prosodic register, producing smaller within-turn distances with zero mutual adaptation. The mixed-effects model includes only a random intercept for session/speaker-pair and the binary same-turn indicator; it has no covariate for prompt identity, topic, or prompt distance. The pattern in Tables 2–3 is consistent with this confound—the effect is largest for Whisper semantic embeddings, exactly the feature space where shared content most directly lowers distance. The absence of significant P2A entrainment does not rescue the P2P findings, since the agent's voice/embedding space differs too much from human speech for the same topic to show as small P2A within-turn distances. This confound is not addressed anywhere in the paper; Section 5's mention of 'unique conversational niches' concerns the agent, not the human-human pairs. Unless prompt effects are removed, the negative γ values do not identify entrainment.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents a new corpus of multiparty conversations between groups of humans (adults or families with children) and a digital agent, and studies entrainment at two timescales: local entrainment within a conversational turn and global entrainment across the session (first vs. last five turns). Using handcrafted acoustic-prosodic features, emotion-model scores/embeddings, Whisper embeddings, and Mimi neural-codec embeddings, the authors fit mixed-effects regressions of pairwise feature distances on a same-turn indicator. They report strong local P2P entrainment in adults across most features, weaker and emotion/deep-feature-only local entrainment in families, no local entrainment with the agent, and limited global entrainment, with children showing global semantic convergence to the agent. The paper interprets these results in terms of rapport, child sensitivity to salient cues, and the special status of the digital agent.","tokens_in":8348,"tokens_out":6449,"duration_ms":69590,"significance":"If the local-entrainment claim is validated, the dataset and the multi-feature analysis pipeline would be a useful addition to multiparty and child-agent interaction research. The study is novel in combining a multiparty human-agent paradigm with both knowledge-driven and foundation-model features, and the use of mixed-effects models with session/speaker-pair random intercepts and Bonferroni correction is methodologically careful. The feature extractors are all externally pretrained and not tuned to the outcome, so the analysis is not circular in the narrow sense. However, the central positive finding—adult local entrainment with other humans—is threatened by a common-cause confound: within-turn human utterances share the same agent prompt, while cross-turn utterances do not. Because the regression in Section 3.1 has no control for prompt identity or topic, the negative γ values in Tables 2–3 may reflect shared content rather than speaker adaptation. This issue is load-bearing for the abstract's main conclusion and needs to be addressed before the result can be trusted.","major_comments":[{"comment":"The local-entrainment model Δ = β_{s,a,b} + γ·I(i,j) + ε has only a same-turn indicator as the fixed predictor, together with a random intercept for session/speaker-pair. By the turn definition in Section 2.1, a turn is one agent prompt followed by participants' responses. Thus, within-turn P2P pairs are simultaneously same-prompt pairs, while cross-turn pairs are different-prompt pairs. A shared prompt can make two responses more similar in topic, lexical content, emotional register, and even prosody with zero mutual adaptation. This confound is not addressed in the paper: Section 5's 'unique conversational niches' sentence concerns the agent–human role separation, not P2P prompt sharing. The pattern in Table 3—largest effects for Whisper semantic embeddings—is exactly what the confound predicts. I request a prompt-controlled reanalysis: add prompt identity/topic as a random or fixed ef","section":"Section 3.1, Eq. (1); Tables 2–3"},{"comment":"The text concludes from the non-significant P2A/C2A local effects that participants 'instinctively treated the agent differently.' This inference goes beyond the data. Null P2A effects cannot serve as a control for the P2P confound because the shared-prompt mechanism does not operate in the same way: an agent's prompt and a participant's response are not acoustically or semantically commensurate, and the agent's synthetic voice occupies a different feature-space region. The absence of a significant P2A effect is also not evidence of absence without a power analysis or equivalence test. I recommend reporting the full P2A/C2A results with effect sizes and uncertainty, and softening the causal interpretation of these null results.","section":"Section 3.2, Tables 2–3"},{"comment":"The global-entrainment analysis compares the first 5 turns with the last 5 turns. Table 1 gives only the median number of turns per session; if any session has fewer than 10 turns, the two windows overlap and the fixed effect is not a clean contrast. The minimum turns per session should be reported, and sessions with fewer than 10 turns should be excluded or handled explicitly. This does not affect the central local-entrainment claim, but it is necessary for the global results to be interpretable.","section":"Section 4.1, Table 1"}],"minor_comments":[{"comment":"Typo: 'prior the the session' should be 'prior to the session.'","section":"Section 3.2"},{"comment":"The caption says '*' denotes p<0.0001 after Bonferroni correction while 'bold values were significant.' Since some entries have significant p-values below 0.05 but above 0.0001 (e.g., Table 3, Emotion dist_cos p=0.0019), the visual encoding is ambiguous. Please clarify the relation between the star threshold, the boldface rule, and the Bonferroni correction.","section":"Tables 2–3 captions"},{"comment":"The notation dist_cos and dist_ℓ2 is used without definition. Define them in the text or in the table captions.","section":"Section 2.2.2"},{"comment":"The paper says the agent's speech was 'automatically removed from the captured audio' but later analyzes participant-to-agent distances. Clarify whether the cached agent audio is the same signal that was removed, and how synchronization with the human channel was ensured.","section":"Section 2.1"},{"comment":"The limitations list mentions the small number of family sessions, the structured interview paradigm, and the interpretability of embeddings, but it does not mention the shared-prompt confound. Since this is the main threat to the central local-entrainment result, it should be acknowledged and addressed.","section":"Section 5"}],"recommendation":"major_revision","confidential_remarks":"The shared-prompt confound is severe enough that the paper should not be accepted in its current form. The central positive result—adult local entrainment—could be an artifact of comparing same-prompt within-turn pairs to different-prompt cross-turn pairs. The authors' Section 5 limitations do not consider this. A prompt-controlled reanalysis is feasible and should be required. If the effect disappears after controlling for prompt identity/topic, the paper's main conclusion would need to be substantially revised; hence major revision rather than reject."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The dataset is a real contribution; the central claim is not yet supported. The local entrainment analysis compares feature distances within a turn against distances across turns, but by the paper's own definition a turn is one agent prompt plus the responses to it. So same-turn utterances share the prompt verbatim, and cross-turn utterances do not. Two answers to the same question will be closer in topic, emotion, and even prosodic register for reasons that have nothing to do with mutual adaptation. The mixed-effects model includes only a same-turn indicator and a random intercept; there is no covariate for prompt identity, topic, or anything that would break that common cause. It is telling that the strongest effects appear in Whisper semantic embeddings, the feature space where shared content most directly lowers distance. The paper's Section 5 discussion of 'unique conversational niches' applies to the agent, not to human-human pairs, so the confound is not addressed.\n\nWhat is genuinely new here is the setting: multiparty groups of adults or families interacting with a digital agent, annotated and segmented by turns. The feature set is broad—pitch, amplitude, emotion, Whisper, Mimi—and the global entrainment analysis uses a different design (first five vs last five turns) that is not subject to the same-prompt confound. The child-agent global entrainment in Whisper features is worth taking seriously, though it rests on ten family sessions.\n\nOther soft spots: the paper drops the non-significant P2A and C2A local modes from the tables instead of reporting them, which is selective reporting. There are no confidence intervals, only p-values. No data or code is released. The family cohort is small.\n\nI would not reject it out of hand. The question is important and the corpus is a contribution. But the headline result—adult local entrainment with other humans—needs a prompt/topic control, effect sizes with uncertainty, and preferably a baseline condition (e.g., the same prompt presented to different groups) before it can be believed. The global results are cleaner and could stand with modest claims. A revision that fixes the local-entrainment analysis would turn this from a promising dataset paper into a solid empirical paper. Worth sending to reviewers, but only with the expectation of substantive revision.","headline":"The multiparty corpus is a real asset; the headline local-entrainment result is most plausibly a shared-prompt artifact.","tokens_in":8861,"tokens_out":3048,"would_cite":false,"duration_ms":30665,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Adult groups entrain strongly with other humans within a turn across acoustic, emotional, and deep-model features, while neither adults nor families entrain with the digital agent locally.","keywords":["multi-party interaction","conversational entrainment","digital agent","child-adult interaction","local entrainment","global entrainment","speech foundation models","acoustic-prosodic features"],"falsifier":"Re-run the local-entrainment regression with a fixed effect for the agent's prompt, or compare only pairs of responses that follow the identical prompt. If the within-turn coefficient for adult pairs drops toward zero, the reported entrainment is a shared-context artifact rather than speaker-to-speaker adaptation.","tokens_in":7951,"feed_emoji":"🗣️","tokens_out":7579,"duration_ms":73517,"temperature":0.7,"pith_summary":"This paper asks whether people adapt their speaking style to each other and to a digital agent in multi-party conversations. It introduces a dataset of adult groups and family groups talking with an animated agent, and measures entrainment at two timescales: within-turn local entrainment and global change over a session. The central finding is that adults strongly entrain to other humans across nearly all measured features, while neither adults nor families locally entrain to the agent. Families entrain mainly on emotional and high-level cues rather than pitch or amplitude, and global entrainment is weak for adults and absent for families. The results imply that digital agents are not treated as conversational peers by adults, and that participant age and interaction structure shape who adapts to whom.","feed_headline":"People entrain with each other, not with the digital agent","feed_subtitle":"Adults align pitch and emotion within a turn; families track emotion only; the agent gets no local adaptation.","key_machinery":"The analysis rests on a mixed-effects regression that compares pairwise distances in a feature between two speakers inside the same turn versus across different turns. The fixed-effect coefficient gamma estimates how much smaller the within-turn distance is after accounting for the speaker pair via a random intercept; a negative gamma is interpreted as local entrainment. Features come from hand-crafted acoustic descriptors (amplitude statistics, pitch, and emotional dimensions) and from pre-trained speech-model embeddings intended to capture semantic and phonetic content. The same regression design is reused for global entrainment by contrasting the first five turns with the last five turns","core_discovery":"The paper's central claim is that in a structured multi-party conversation with a digital agent, local entrainment—speakers becoming more similar within a single turn—is consistent among human speakers but does not extend to the agent. For adult pairs, almost every hand-crafted and model-based feature shows a significant negative within-turn coefficient, meaning adult participants align pitch, amplitude, emotion, and semantic/phonetic similarity with other humans. Family groups show a narrower pattern: entrainment on emotional arousal, dominance, and deep embeddings, but not on pitch or amplitude; child–guardian pairs follow the same pattern. Neither cohort shows local entrainment with the a","pith_inferences":["A caution: the local-entrainment design compares within-turn responses to cross-turn responses, so it cannot rule out that two speakers are independently aligning to the same agent prompt rather than to each other. Adding prompt identity as a covariate, or restricting comparisons to identical prompts, would test this directly.","If that common-cause explanation survives, the family emotional-entrainment result—found with the same design—would also need re-examination, since both speakers might be reacting to the same question.","The cohort difference suggests a testable extension: children may track salient emotional cues while ignoring low-level acoustic detail; a playback study that manipulates only pitch or only emotional tone could separate which cues children actually follow.","For agent design, the results imply that a voice agent wanting rapport with adult groups may need to actively entrain to the humans rather than wait for adaptation; the child-only global effect hints that relationship-building with younger users requires a longer arc."],"forward_implications":["If adults consistently entrain with each other but not with the agent, digital agents in multi-party settings should not expect human speakers to accommodate to the agent's acoustic or emotional style.","Entrainment is feature-specific: family groups converge on emotional arousal and dominance while showing no pitch or amplitude convergence, so a single global entrainment score would hide the effect.","The absence of local agent entrainment in both cohorts suggests the agent occupies a distinct conversational role rather than a peer role in question-driven interactions.","Children's global semantic convergence with the agent, despite no local entrainment, indicates that child–agent adaptation unfolds over longer timescales and may reflect relationship-building.","The apparent amplitude decoupling between children and the agent reflects children becoming louder and more variable over the session, not a failure to entrain."],"fun_headline_variants":["Humans sync speech with each other, not bots","Digital agents don't get speech entrainment","In group chats, humans align, but not with the agent","Speech entrainment is human-to-human, not human-to-agent","People sync with people, not with the digital agent"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The local-entrainment results assume that two human speakers sounding alike within a turn is caused by mutual adaptation, not by both of them responding to the same agent prompt.","fun_headline_variants_meta":{"raw":{"variants":["Humans sync speech with each other, not bots","Digital agents don't get speech entrainment","In group chats, humans align, but not with the agent","Speech entrainment is human-to-human, not human-to-agent","People sync with people, not with the digital agent"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000955,"raw_usage":{"total_tokens":3869,"prompt_tokens":668,"completion_tokens":3201,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":412,"completion_tokens_details":{"reasoning_tokens":3131}},"tokens_in":412,"tokens_out":3201,"duration_ms":20801,"temperature":1.0,"reasoning_tokens":3131,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T04:05:09.803416+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the local-entrainment regression with a fixed effect for the agent's prompt, or compare only pairs of responses that follow the identical prompt. If the within-turn coefficient for adult pairs drops toward zero, the reported entrainment is a shared-context artifact rather than speaker-to-speaker adaptation.","supporting_citations":[],"review_version":1}