{"id":"054f625f-4f8a-4c44-9ccf-51babda92677","arxiv_id":"2508.02376","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A photorealistic speaking avatar elicits longer, more information-dense survey responses than a text chatbot, but with worse clarity and no satisfaction benefit.","lead":"What happened: 80 UK participants answered psychometric surveys either by talking to a photorealistic AI avatar or typing to a chatbot. The avatar made responses longer and more information-dense, but slightly less clear and relevant, with no gain in satisfaction. For survey researchers, this is a proof of concept that embodied AI interviewers could reduce satisficing in unmoderated online research, though technical glitches and the Uncanny Valley still limit adoption.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Claimed informativeness advantage for the embodied agent is indistinguishable from a speaking-vs-typing channel effect: the surprisal-sum measure in §4.6 is not modality-invariant, and the paper's own clarity/relevance results show the channel differs.","rationale":"Both the reader and I identify the same load-bearing assumption: the informativeness measure is not modality-invariant, so the core RQ1 result cannot be cleanly attributed to embodiment. The surrounding evidence strengthens this concern. The content-based measures (specificity null; clarity and relevance favoring text) point in the opposite direction, and the CFQ sub-analysis fails to replicate. The time-efficiency claim is more defensible because it directly compares speaking versus typing time for the same task, but the response-quality claim is the scientifically central one. None of this impugns the engineering value of the paper; the instrument, the between-subjects design, and the released repository are real contributions. It does mean the abstract's 'more informative, detailed responses' overstates what the data can establish. Since the reader already set CONDITIONAL, I recommend no change to that verdict; the revisions should include a normalized-informativeness sensitivity analysis, explicit discussion of the channel confound, and a more circumscribed conclusion.","tokens_in":26047,"tokens_out":5633,"duration_ms":66370,"concrete_test":"Recompute §4.6 informativeness on the released repository data as per-token mean surprisal after removing disfluencies and non-content tokens, rather than as summed surprisal. If the embodied-versus-text difference on this normalized measure becomes non-significant or drops below r≈.2, the reported r=.39 is a volume/channel artifact and the 'more informative' claim must be restricted to response volume. This uses only the existing dataset and requires no new data collection.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim that embodied agents elicit more informative responses rests on the 'informativeness' measure (§4.6): the sum of per-word surprisal from the wordfreq database. For that measure to support an embodiment effect, it must be invariant across communication channels. It is not. Spoken transcripts and typed text differ systematically in vocabulary register, disfluencies, and ASR artifacts; speech recognition errors can produce rare, high-surprisal tokens that inflate the summed score without adding substantive content. The embodied arm also produced more words (M=29.32 vs 14.3, r=.40, §5.2), so the summed surprisal would rise even if the underlying content were equivalent. The paper's own §5.1 results make the channel confound visible: embodied responses were much less clear (r=.63) and slightly less relevant (r=.26), exactly the pattern expected from transcribed spontaneous speech. The design also bundles voice input, video avatar, and turn-taking behavior in the embodied arm (§3.1), so even the word-count difference cannot be attributed specifically to embodiment. §5.4 shows the effect is fragile: within the CFQ sub-survey, informativeness and word count are non-significant (p=.081, p=.067) despite the same modality difference. If the informativeness advantage does not survive content-normalization, the headline conclusion reduces to 'people speak more when talking rather than typing,' not 'ECAs improve response quality.'","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents the Virtual Agent Interviewer (VAI), a survey instrument that embeds a photorealistic embodied conversational agent (ECA) with voice interaction, and compares it against a text-based chatbot in a between-subjects experiment with 80 UK participants across two psychometric surveys (BFI-2-S and CFQ). The authors report that the embodied agent produces significantly higher word counts and informativeness, shorter participant responding time, and no difference in satisfaction, and they interpret these results as evidence that ECAs can improve response quality and engagement. The paper includes a mixed-methods analysis with qualitative feedback, and it makes the instrument, code, and data publicly available. The central claim, however, is that embodiment itself drives the improvement, whereas the experimental design bundles voice input, a video avatar, and turn-taking into one condition that is compared to a typing-only chatbot.","tokens_in":26273,"tokens_out":4204,"duration_ms":57457,"significance":"If the causal interpretation were supported, this would be a useful proof-of-concept for bringing ECA-mediated moderation to online surveys, with practical implications for reducing satisficing and improving open-ended response quality. The paper is transparent about many null results and limitations, and the open repository and the detailed description of the instrument are commendable. However, the study's headline inference is not currently supported because the 'informativeness' measure is confounded with response length and with the speech-versus-text channel, and the effect is not robust across the two surveys. These issues reduce the significance of the current findings, although the methodological contribution may still be useful to the community.","major_comments":[{"comment":"The 'informativeness' measure is defined as the sum of per-word surprisal from wordfreq (§4.6). This sum scales with word count, and §5.2 reports that embodied-agent responses contain roughly twice as many words (M=29.32 vs 14.3, r=.40). Consequently, the reported informativeness advantage (r=.39) could simply reflect that participants produce more words when speaking than when typing; it does not establish that the same underlying content is more informative. The paper's own results show that the two channels differ: embodied responses were less clear (r=.63) and slightly less relevant (r=.26), consistent with a speaking-versus-typing channel effect. The authors should re-analyze the data using an information-density measure (e.g., per-word surprisal, or surprisal per 100 characters), or control for word count as a covariate, before claiming an informativeness advantage. A concrete test would be to compare the embodied condition with a voice-only condition that does not have a visible avatar.","section":"§4.6, §5.1, §5.2"},{"comment":"The headline effects on informativeness and word count are not robust across the two surveys: within CFQ, the comparisons are non-significant (informativeness p=.081, word count p=.067), while only BFI-2-S shows significant effects (r=.47). This pattern suggests that the overall significant result may be driven by only one of the two questionnaires, and it limits the abstract's generalization that embodied agents 'contribute significantly to more informative, detailed responses' across surveys. The authors should report a formal survey-by-condition interaction test, and should temper the abstract and conclusions accordingly.","section":"§5.4"},{"comment":"The embodiment manipulation is compounded with response modality: the embodied condition uses voice input plus a video avatar, while the baseline uses typed text input. Thus even the robust word-count difference can be explained by 'people speak more than they type' rather than by the embodied conversational agent per se. The paper partly acknowledges this in the clarity discussion, but it does not treat the modality confound as a threat to the central 'embodiment improves quality' claim. A comparison with an additional condition (e.g., a voice-only non-visual agent, or a text-input condition with an avatar) would be necessary to isolate the contribution of embodiment.","section":"§3.1, §4.4"},{"comment":"The authors report a large number of statistical tests without any correction for multiple comparisons, and the stated RQ framework includes four research questions with several measures each. While the strongest effects (p<.001) would likely survive conservative correction, the overall familywise error rate is not addressed. Reporting corrected p-values or a pre-specified analysis plan would strengthen the reliability of the secondary findings such as self-disclosure and the survey-specific results.","section":"§5.1, §5.2, §6.2"}],"minor_comments":[{"comment":"The sentence 'using on AI-driven video generation' contains a grammatical error; it should be 'using AI-driven video generation.'","section":"Abstract"},{"comment":"The sentence 'Text-based variants' answers tend to be longer' under Figure 9 contradicts the surrounding text, which states that the embodied agent produced significantly more words; this should be corrected to 'Embodied variants' answers tend to be longer.'","section":"§5.2"},{"comment":"The word 'Colloquilally' is a typo for 'Colloquially.'","section":"§2.2"},{"comment":"The data-labeling procedure uses GPT-4.1 with majority voting, and the authors report high correlations with manual labels; however, they do not state whether the manual reviewers were blind to the experimental condition, which would be relevant given the large clarity differences between conditions.","section":"§4.6"},{"comment":"In Table 1 and the surrounding text, there are occasional spacing anomalies (e.g., 'T able' at the top of Table 1); these should be cleaned up for publication.","section":"§5.3"}],"recommendation":"major_revision","confidential_remarks":"The central issue is that the paper's main claim is not supported by the design: the embodied condition differs from the baseline in both modality and embodiment, and the 'informativeness' measure is essentially a word-count measure. The authors should be encouraged to re-analyze with per-word or density measures, and to either add a control condition or substantially soften the causal claims. The contribution of the instrument and the open data is real, but the conclusions as currently written overstate the evidence."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: the paper is a solid, readable proof-of-concept with transparent statistics and genuinely useful open artifacts, but its headline is one size too big. The design supports 'people say more, faster, when speaking to a talking avatar than when typing to a chatbot'; it does not support the stronger claim that embodiment improves response quality.\n\nWhat is new: this is the first comparison I know of that pairs a photorealistic, LLM-driven speech avatar with a text chatbot for open-ended survey follow-ups on a general-population sample. The between-subjects design is clean, the power analysis is reported, the two questionnaires give some cross-validation, and the authors report nulls and subgroup differences instead of hiding them. The qualitative material on turn-taking delays and Uncanny Valley reactions is useful for anyone building this kind of tool. Credit also for shipping data and code.\n\nThe soft spots are real and central. The 'informativeness' measure is a sum of per-word surprisal from wordfreq. That is not modality-invariant. Spoken transcripts differ from typed text in register, disfluency, and ASR artifacts, and the embodied arm produced roughly twice as many words, so the summed surprisal would rise even if the underlying content were identical. The paper's own results make the channel confound visible: embodied responses were substantially less clear (r=.63) and slightly less relevant (r=.26). The design also bundles voice input, video avatar, and turn-taking behavior, so the word-count advantage cannot be attributed specifically to embodiment. And in the CFQ survey the word-count and informativeness effects are not significant (p=.067, p=.081), which undercuts the general claim. The abstract's 'higher yet more time-efficient engagement' also overstates things: total completion time was not faster, and satisfaction did not differ.\n\nNone of this sinks the paper. The instrument works, the data are honestly presented, and the authors' own limitations section shows they know the boundaries. But the abstract and conclusion should be rewritten to say what was actually found: a talking avatar elicits longer and faster verbal responses than a text chatbot, with no measurable satisfaction cost, and with some evidence of lower clarity and relevance. A sensitivity analysis for informativeness (e.g., normalizing by word count, or comparing only clean transcripts) would address the main confound.\n\nWho is this for: HCI and survey-methodology researchers working on conversational agents or unmoderated research. It deserves a serious referee, but with the expectation of major revision. I would not desk-reject it.","headline":"A solid, transparent proof-of-concept showing people talk more and faster to a speech avatar than a text chatbot, but the informativeness claim is confounded with speaking-versus-typing and the abstract oversells it.","tokens_in":26898,"tokens_out":3057,"would_cite":true,"duration_ms":38389,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Talking to a photorealistic avatar yields longer, more informative survey answers than typing to a chatbot, in less speaking time.","keywords":["conversational agent","AI-mediated communication","photorealistic avatar","user engagement and satisfaction","virtual human","chatbot","survey response quality","open-ended questions"],"falsifier":"Have the same participants answer the same open-ended questions under three conditions—typing to a chatbot, speaking to a voice-only agent, and speaking to the animated avatar—and compare word-surprisal informativeness and word counts; if voice-only speech matches the avatar, the central effect is driven by speech rather than embodiment, and the paper's claim would need to be narrowed.","tokens_in":25773,"feed_emoji":"🗣️","tokens_out":18168,"duration_ms":177202,"temperature":0.7,"pith_summary":"This paper argues that putting a photorealistic, voice-enabled avatar into an online survey changes what people say: participants who talked to the avatar gave roughly twice as many words and twice as much information per open-ended answer as participants who typed to a text chatbot, while spending significantly less time speaking their answers. The authors built a working instrument, the Virtual Agent Interviewer, that combines AI-driven video generation, speech recognition, and a large language model, and tested it on 80 UK panel participants completing two psychometric surveys. They frame the finding as a step toward closing the gap between unmoderated online surveys and human-moderated interviews, where personal engagement and perceived accountability are higher. They also report that satisfaction ratings did not differ between the two agents, with qualitative feedback pointing to turn-taking delays and Uncanny Valley reactions as the main sources of friction.","feed_headline":"Talking survey avatars draw richer answers than chatbots—in less time","feed_subtitle":"An 80-person trial found the embodied agent about doubled word count and informativeness of open-ended answers.","key_machinery":"The central object is the Virtual Agent Interviewer (VAI), a prototype survey tool that embeds a conversational agent into the survey flow. In the embodied condition the agent is a photorealistic animated video avatar that speaks; in the text condition the same conversation logic runs as a typed chatbot, so the design isolates the addition of embodiment plus speech against the chatbot baseline. The response-quality machinery is a set of measures adapted from conversational-survey research: informativeness as summed word surprisal, that is, the inverse frequency of the response's words, from the wordfreq database, plus human- and LLM-coded specificity, relevance, clarity, self-disclosure, and sentiment. The argument turns on comparing these measures across the two conditions together with temporal engagement metrics such as responding time and transition time.","core_discovery":"The paper's central claim is that embodied conversational agents can improve the quality and engagement of survey responses without lowering satisfaction. In a between-subjects experiment that held conversation logic constant, participants who spoke to an animated video avatar produced more informative responses ($M=285.41$ vs $M=142.23$, $U(80)=1164$, $z=3.50$, $p<.001$, $r=.39$), used more words ($M=29.32$ vs $M=14.3$, $U(80)=1171$, $z=3.57$, $p<.001$, $r=.40$), and spent less time speaking (responding time $U(80)=531$, $z=2.59$, $p=.010$, $r=.29$) than participants who typed to a text-based chatbot. The authors interpret the steeper time-informativeness slope for the avatar as evidence that speaking to a humanlike agent is a more efficient communication channel, not merely a longer one. They note that spoken answers scored lower on clarity ($r=.63$) and relevance ($r=.26$), which they attribute to the spontaneous, unedited character of speech rather than to worse content. Satisfaction differences were not statistically significant, and the paper presents this as a favorable baseline for embodied agents given the prototype's technical limitations.","pith_inferences":["A sharper test of embodiment itself would compare the avatar against a voice-only agent with no face; the reported design cannot separate the effect of speech modality from the effect of a humanlike body and face.","The word-surprisal measure may overstate the informativeness advantage if natural speech systematically draws on rarer vocabulary for the same underlying content; re-analyzing with frequency-matched or lemmatized transcripts would check this.","If the pattern replicates, a practical design rule follows: reserve embodied agents for broad, identity-relevant open-ended questions and expect smaller gains for narrowly factual or more sensitive items.","A direct extension would test whether telling participants that a human moderator is watching the avatar changes response length, separating perceived accountability from the visual design itself."],"forward_implications":["Unmoderated survey tools could collect more detailed qualitative data from open-ended questions without asking participants to spend more time producing each answer.","Because the time-informativeness slope was steeper for the avatar, speaking to an embodied agent appears to transmit more information per second than typing to a chatbot.","Satisfaction being statistically unchanged suggests that adding a photorealistic face to a survey agent does not carry an inherent satisfaction penalty, though preference data show about half of embodied-condition participants would still switch to text.","The larger and significant effects in the personality survey compared with the marginal effects in the cognitive-failures survey imply that the benefit of embodied agents may depend on question breadth and sensitivity.","Lower clarity and relevance in spoken responses mean embodied-agent transcripts may require more preprocessing or a visible question prompt to keep answers on target."],"supporting_citations":[{"why":"It is the closest prior ECA study; its limitations motivate the paper's LLM-driven design and general-population sample.","marker":"Zhu and Broadbent (2025)"},{"why":"It supplies the conversational-survey method and the response-quality measures (informativeness, specificity, relevance, clarity) that the paper adapts.","marker":"Xiao et al. (2020)"},{"why":"It establishes the chatbot-survey baseline the paper extends, including the finding that conversational style can reduce satisficing.","marker":"Kim et al. (2019)"},{"why":"It defines careless responding as a threat to online survey quality, the problem the embodied agent is meant to address.","marker":"Ward and Meade (2018)"},{"why":"It supplies the wordfreq database from which informativeness (summed word surprisal) is computed.","marker":"Speer (2022)"},{"why":"It is the source of the BFI-2-S personality inventory used as one of the two survey instruments.","marker":"Soto and John (2017)"},{"why":"It is the source of the Cognitive Failures Questionnaire used as the second survey instrument.","marker":"Rast et al. (2009)"},{"why":"It shows that LLM follow-up questions can elicit more detail but also repetition, informing the agent's conversation-module design.","marker":"Kuric et al. (2024)"}],"fun_headline_variants":["Survey avatars beat chatbots on detail and speed","Talking to a video avatar yields richer survey answers","Avatar surveys: more words, better info, less time","Embodied agents double survey response detail","Avatar vs chatbot: richer open answers in less time"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that word-surprisal—a count of how rare a word is in English—measures informativeness equally for speech and for typing; if people naturally use rarer words when speaking, the avatar's informativeness advantage could be an artifact of comparing two different communication channels.","fun_headline_variants_meta":{"raw":{"variants":["Survey avatars beat chatbots on detail and speed","Talking to a video avatar yields richer survey answers","Avatar surveys: more words, better info, less time","Embodied agents double survey response detail","Avatar vs chatbot: richer open answers in less time"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00083,"raw_usage":{"total_tokens":3684,"prompt_tokens":1061,"completion_tokens":2623,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":677,"completion_tokens_details":{"reasoning_tokens":2550}},"tokens_in":677,"tokens_out":2623,"duration_ms":23505,"temperature":1.0,"reasoning_tokens":2550,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T04:57:40.928992+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have the same participants answer the same open-ended questions under three conditions—typing to a chatbot, speaking to a voice-only agent, and speaking to the animated avatar—and compare word-surprisal informativeness and word counts; if voice-only speech matches the avatar, the central effect is driven by speech rather than embodiment, and the paper's claim would need to be narrowed.","supporting_citations":[{"cited_title":"Factor structure and measurement invariance of the cognitive failures questionnaire across the adult life span","cited_arxiv_id":null,"evidence_quote":"It is the source of the Cognitive Failures Questionnaire used as the second survey instrument."}],"review_version":1}