{"id":"4632a845-4361-44c4-8a6c-9d8b9f53b114","arxiv_id":"2411.17157","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"For three of four role-playing chatbots, perceived anthropomorphism correlated positively with users' media dependency, and satisfaction mediated most of the relationship.","lead":"This study of a Chinese role-playing chatbot community found that the more human users perceived a chatbot to be, the more they tended to depend on it, with satisfaction an important bridge. The result held for three of four characters tested, and interviews explained when and why the pattern breaks.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The four roles differ in character lore and prior familiarity, so the reported correlations may reflect attachment to the original characters rather than anthropomorphism; the authors' own Role 2 interviews show this confound can dominate.","rationale":"The paper's central claim is that perceived anthropomorphism increases media dependency, with satisfaction as a mediator. For this claim to hold, differences in dependency across the four roles must be attributable to the anthropomorphism construct rather than to any of the many other ways the roles differ. The selection procedure in Section 3.2 deliberately chose roles spanning different anthropomorphism scores, but those roles are also different anime/game characters with distinct backstories, fan communities, and prior exposure patterns. The authors' own interview data for Role 2 demonstrate that when a character has a complex worldview and users have varying prior knowledge, dependency ratings can be driven by familiarity and preference rather than humanness. If this pattern generalizes to the other three roles, the observed correlations in Tables 2, 4, and 5 are confounded and the causal framing in the abstract and conclusion is unsupported. The cross-sectional design compounds this problem because it cannot rule out reverse causality or third variables such as general enjoyment of the source franchise. I agree with the reader's weakest_assumption that this is the core threat to the central claim. The paper has real strengths—reported reliability, clear tables, and honest discussion of Role 2's failure—but the central causal claim depends on isolating anthropomorphism from character identity, which the current design does not do. The verdict should remain CONDITIONAL, with the condition being that the authors either obtain experimental evidence that manipulates anthropomorphism while holding character constant, or substantially soften their causal language and acknowledge that the correlations may reflect character attachment rather than anthropomorphism. A controlled follow-up study, as described in the concrete test, would settle whether the concern actually lands.","tokens_in":19287,"tokens_out":4413,"duration_ms":45706,"concrete_test":"Run a between-subjects experiment holding character identity constant: randomly assign participants to interact with the same chatbot (e.g., Hu Tao) scripted in two styles—high-anthropomorphic (emotional expressions, personal revelations, human-like fluency) versus low-anthropomorphic (functional, generic, minimal social cues)—for the same 10-day period. Measure perceived anthropomorphism (HRIES), satisfaction, and media dependency. Compute the indirect effect of condition on dependency via satisfaction with a bootstrap 95% CI. If the CI includes zero or the total effect is not significant, the original role-level correlations are likely confounded with character familiarity/narrative appeal.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"Section 3.2 selects four chatbots via stratified sampling of perceived anthropomorphism, but these roles also differ systematically in character background, game/anime lore, and likely user familiarity. The interviews in Section 4.3.1 (R1.1–R1.3) show that for Hu Tao (Role 2), prior knowledge, expectations, and personal preference drove dependency ratings, not perceived humanness. If similar dynamics hold for Roles 1, 3, and 4, the significant regressions in Tables 2, 4, and 5 could reflect users' pre-existing attachment to the source characters or narrative appeal rather than the anthropomorphism construct measured by HRIES. The design is a single cross-sectional questionnaire after 10 days of interaction (Section 3.3), so it cannot establish temporal precedence or rule out third variables; no measure of prior familiarity, fandom, or character liking was collected. Because the central claim is that anthropomorphism increases media dependency through satisfaction, this confounding directly threatens the causal interpretation, not just the mediation sub-path.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript reports a mixed-method user study of the Chinese role-playing chatbot platform Xuanhe AI. After a preliminary survey selected four chatbots spanning low to high perceived anthropomorphism (Asuka Langley Soryu, Hu Tao, Yandere Girlfriend, Satoru Gojo), 149 users were recruited to interact with all four chatbots for ten days, and 108 valid questionnaires measured perceived anthropomorphism (HRIES), media dependency, and satisfaction. Hierarchical regressions for each role tested whether anthropomorphism predicts dependency and whether satisfaction mediates; for Roles 1, 3, and 4 the direct and mediated paths were significant, while for Role 2 the direct path was significant but the satisfaction path was not (B = -0.017, p = .946). Semi-structured interviews with ten deviant-case users identified prior knowledge and preferences, real-life distractions, and conscious self-control as factors interfering with the hypothesized relationship.","tokens_in":19473,"tokens_out":5353,"duration_ms":47374,"significance":"If the association were established, the paper would extend media dependency theory to LLM-based role-playing chatbots in a Chinese context and offer concrete design implications for anthropomorphic chatbot design and anti-addiction features. The study has notable strengths: the hypothesis is stated before the quantitative analysis in Section 3.2, reliability statistics are reported per role (Table 1), the null result for Role 2 is reported rather than hidden, and the qualitative follow-up is transparently exploratory. The main value is empirical and contextual. However, the contribution is currently weakened by causal overreach and by confounds between anthropomorphism and the specific characters chosen, so the significance depends on the revisions described below.","major_comments":[{"comment":"The hypothesis is stated causally ('will increase') and the abstract reports a 'significant positive correlation' with 'satisfaction mediating', but the design is a single cross-sectional questionnaire administered after ten days of interaction (Section 3.3), with no baseline measure and no manipulation of anthropomorphism. The data therefore support only a correlational claim; the causal wording should be removed or explicitly labeled as theoretical motivation, and the mediation path should be described as consistent with the proposed model rather than as evidence of mechanism.","section":"Abstract and Section 3.2"},{"comment":"The four chatbots were selected to vary in perceived anthropomorphism, but they also differ systematically in franchise, lore, and likely prior familiarity (Asuka, Hu Tao, Yandere Girlfriend, Satoru Gojo). The interviews show for Role 2 that non-gamers' unfamiliarity (R1.1), lore-based expectations (R1.2), and preferences for other characters (R1.3) drove dependency ratings, not perceived humanness. No measure of prior familiarity, fandom, or character liking was collected. Because the same confound could operate for Roles 1, 3, and 4, the regressions in Tables 2, 4, and 5 do not isolate anthropomorphism as the active ingredient; add a control for prior character familiarity/liking or reanalyze restricted to users with comparable familiarity, and temper the causal interpretation accordingly.","section":"Section 3.2 and Section 4.3.1"},{"comment":"The mediation analyses use Baron and Kenny's causal-steps approach and do not report a confidence interval or bootstrap test for the indirect effect; with Role 2 the satisfaction coefficient in Model 3 is B = -0.017, p = .946, yet the paper still interprets Role 2 as an unsupported mediation case based on path significance. More importantly, the same 108 participants rated all four chatbots, so the four role-level regressions are not independent; ignoring within-subject correlation can deflate standard errors and inflate the reported p-values. I recommend multilevel or repeated-measures analysis with role as a within-subject factor, and bootstrap or Monte Carlo confidence intervals for the indirect effects.","section":"Section 3.4"},{"comment":"The deviant-case interviews are described as explaining why Role 2 failed, but the selection of ten users with 'significant deviations' is post hoc and the criteria for identifying those deviations are not defined; the grounded-theory categories (R1-R3) are therefore exploratory hypotheses, not confirmatory evidence. This is acceptable as interpretation, but the paper should state that these categories were generated after seeing the regression results and cannot themselves validate the proposed mechanism.","section":"Section 4.1 and Section 4.2"}],"minor_comments":[{"comment":"Section 3.3 reports the study ran from January 19 to January 29, 2024, but Section 4.3.2 states the user study was conducted in February; please correct the inconsistency.","section":"Section 3.3 and Section 4.3.2"},{"comment":"The limitation statement says that half of the selected chatbots came from the 'anime characters' channel, but three of the four roles (Asuka, Hu Tao, Satoru Gojo) are anime/game/manga-derived; adjust the statement.","section":"Section 5.4"},{"comment":"The grounded-theory coding description does not report the number of coders or inter-coder agreement; please add this information to support the trustworthiness of the qualitative analysis.","section":"Section 4.2"},{"comment":"The abstract says 149 users were invited but does not mention that the analyses are based on 108 valid participants; please include the valid sample size or qualify the abstract accordingly.","section":"Abstract and Section 3.3"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is readable and the mixed-method design is appropriate for a conference paper, and the transparent reporting of Role 2's null result is a credit. My main concern is that the causal framing and the role-level confounds need substantive rather than cosmetic revision: the abstract should be reworded to correlational language, sensitivity analyses or at least an explicit discussion of the familiarity confound should be added, and the within-subject structure of the data should be accommodated in the statistical model. I do not see this as a rejectable paper, but the current version would overstate what the evidence supports if published without these changes."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: honest, modest mixed-methods study of anthropomorphism and media dependency on a Chinese role-playing chatbot platform. The headline result—positive correlation with satisfaction as mediator—holds for three of four roles, and the authors report the null rather than hiding it. The abstract oversells causality, and the character-lore confound is real. Still, this deserves a serious referee.\n\nWhat's new: they apply media dependency theory, usually used for traditional media, to LLM role-playing chatbots on Xuanhe AI. The four-role stratified selection and 10-day interaction window are reasonable attempts at naturalistic data. Their best move is what they do with the null: instead of dropping Role 2, they interview ten deviant-case users and identify three factors—prior knowledge and preferences, real-life distraction, conscious self-control—that interrupt the anthropomorphism-dependency link. That is genuinely useful for designers and for future work on emotional AI attachments.\n\nSoft spots, in proportion. First, the causal language in the abstract ('will increase') overstates a cross-sectional, one-time survey. There's no baseline and no manipulation; they can claim association, not causation. Second, the four roles differ in game/anime lore and user familiarity, not just perceived humanness. The authors' own interviews show that for Hu Tao, prior knowledge and personal preference drove dependency ratings—so the confound is demonstrated in their data, not hypothetical. The same dynamic likely operates for the other roles. Third, the mediation analysis is Baron-Kenny without bootstrap confidence intervals and treats each role as an independent sample, even though the same 108 participants rated all four roles; that ignores within-subject correlation and may inflate significance. Minor: the exclusion criteria for the 108 are vague and the interview selection is underdescribed.\n\nWho this is for: HCI researchers studying companion chatbots and design practitioners. If you read the abstract you'll be misled; if you read it as an exploratory correlation with a thoughtful qualitative coda, it's a legitimate data point. I'd send it to peer review, asking for correlational framing and attention to the confounds.","headline":"A modest, honest mixed-methods study; the anthropomorphism-dependency correlation holds for three of four roles, but the abstract's causal claim and the character-lore confound need to be addressed.","tokens_in":19959,"tokens_out":3690,"would_cite":true,"duration_ms":32954,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"More human-like role-playing chatbots create stronger media dependency, with satisfaction carrying the effect for three of four tested roles.","keywords":["Chatbot","Anthropomorphism","Media Dependency","Uses and Gratifications","Human-Machine Communication","role-playing chatbots","mediation analysis","large language models"],"falsifier":"Re-run the study with the same four characters but statistically control each user's prior familiarity with the character's source material, or present users with two versions of one character that differ only in linguistic human-likeness; if the anthropomorphism–dependency coefficient vanishes or reverses under those controls, the paper's central claim fails.","tokens_in":19041,"feed_emoji":"🤖","tokens_out":8805,"duration_ms":74535,"temperature":0.7,"pith_summary":"This paper tests a chain of influence for role-playing AI chatbots: the more human-like a chatbot seems, the more satisfied users become, and the more satisfied users become, the more they depend on it. The authors recruited 149 users of the Chinese platform Xuanhe AI, had them interact with four popular chatbots for ten days, and analyzed questionnaire responses from 108 valid participants using an anthropomorphism scale, a satisfaction scale, and an adapted media-dependency scale. They found the predicted correlation and satisfaction-mediated path for three of the four roles; for the fourth, perceived humanness still predicted dependency but satisfaction did not mediate it. Follow-up interviews attribute that exception to users' prior knowledge of the character, real-life distractions, and deliberate self-control. If the claim generalizes, it gives designers and researchers a concrete mechanism through which making a chatbot more human-like can deepen emotional attachment, and a set of factors that can interrupt that attachment.","feed_headline":"Humanlike chatbots breed dependency, satisfaction is the bridge","feed_subtitle":"In a ten-day user study on a Chinese AI role-play platform, perceived humanness predicted dependency in three of four AI roles.","key_machinery":"The load-bearing structure is a three-variable mediation model built on media dependency theory: perceived anthropomorphism (X) to user satisfaction (M) to media dependency (Y), with demographic controls. Anthropomorphism is measured with the HRIES scale, a 16-item instrument covering sociability, agency, animacy, and disturbance; dependency is measured with a six-item scale adapted from the Facebook Addiction Scale; satisfaction is measured with a three-item scale. The argument is carried by estimating this mediation model separately for each of the four chatbots and then using a grounded-theory analysis of ten interviews to explain the one role where the mediation path disappeared.","core_discovery":"The central claim is that perceived anthropomorphism in role-playing chatbots is positively associated with users' media dependency, and that user satisfaction is a genuine mediator of that association. In the regression models for Roles 1, 3, and 4, anthropomorphism significantly predicted dependency and satisfaction, and satisfaction remained a significant positive predictor of dependency when both were entered together, supporting the hypothesized chain. For Role 2, the direct path from anthropomorphism to dependency was significant but the satisfaction path was not, so the mediation hypothesis was not supported for that role; the authors use interview data to show that character familiarity, expectations from prior knowledge, life circumstances, and deliberate emotional self-control can break the chain.","pith_inferences":["A natural next test would compare two versions of the same character that differ only in perceived humanness, isolating anthropomorphism from character lore and prior familiarity; this would be the cleanest way to confirm the causal direction.","The Hu Tao exception implies that familiarity with a character can cut both ways—it can deepen engagement for fans, yet raise expectations that make a chatbot's errors more disappointing—which is a design tension worth testing directly.","If the satisfaction-mediated path is real, platforms aiming to reduce compulsive use might intervene on satisfaction, for example by reducing emotional reward or adding friction, rather than by removing human-like features altogether.","The current data cover moderate levels of anthropomorphism only, so an open question is whether very high levels eventually reverse the positive effect through the uncanny valley; the paper does not reach that range."],"forward_implications":["Design choices that raise a chatbot's perceived humanness are also choices that can raise users' dependency, so anthropomorphic features carry an attachment cost as well as an engagement benefit.","Satisfaction is the main conveyor of that effect for most roles, but not all; when users bring strong prior knowledge or preferences about a character, satisfaction may be disconnected from dependency.","The same pattern appearing across anime, game, and meme-derived roles suggests the effect is not tied to one genre of chatbot content.","Since real-life distractions and conscious self-control weakened dependency in the interviews, the relationship is conditional rather than automatic and can be moderated by the user's situation."],"supporting_citations":[{"why":"Supplies the HRIES scale, the 16-item instrument used to measure perceived anthropomorphism of each chatbot.","marker":"[49]"},{"why":"Provides the definition of anthropomorphism that motivates the hypothesis.","marker":"[16]"},{"why":"Establishes the precedent of connecting usage motivations to media dependency, applied here to Facebook.","marker":"[17]"},{"why":"Extends media dependency thinking to AI companionship, the direct theoretical background for dependency on chatbots.","marker":"[57]"},{"why":"Offers the relationship-development pattern the authors use to interpret why satisfaction builds into dependency.","marker":"[48]"},{"why":"Provides the grounded-theory method used to code interview responses and derive the three interfering factors.","marker":"[36]"},{"why":"Sets the Cronbach's alpha acceptance threshold used to justify scale reliability.","marker":"[52]"},{"why":"Shows media dependency theory applied to an AI-driven app, grounding the model's satisfaction-to-attitude-to-behavior chain.","marker":"[8]"}],"fun_headline_variants":["Humanlike AI chatbots breed dependency, satisfaction mediates","When AI feels human, users grow dependent — satisfaction is key","Anthropomorphism in role-playing chatbots drives dependency, study shows","Role-playing AI: Perceived humanness boosts dependency via satisfaction","AI chat personas: More human, more dependent, satisfaction as bridge"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The four chosen chatbots are treated as varying mainly in perceived humanness, but they also differ in fame, backstory, and users' prior familiarity, so the measured dependency could be caused by character attachment rather than by anthropomorphism.","fun_headline_variants_meta":{"raw":{"variants":["Humanlike AI chatbots breed dependency, satisfaction mediates","When AI feels human, users grow dependent — satisfaction is key","Anthropomorphism in role-playing chatbots drives dependency, study shows","Role-playing AI: Perceived humanness boosts dependency via satisfaction","AI chat personas: More human, more dependent, satisfaction as bridge"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000167,"raw_usage":{"total_tokens":1236,"prompt_tokens":905,"completion_tokens":331,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":521,"completion_tokens_details":{"reasoning_tokens":246}},"tokens_in":521,"tokens_out":331,"duration_ms":4835,"temperature":1.0,"reasoning_tokens":246,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T12:26:33.965963+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the study with the same four characters but statistically control each user's prior familiarity with the character's source material, or present users with two versions of one character that differ only in linguistic human-likeness; if the anthropomorphism–dependency coefficient vanishes or reverses under those controls, the paper's central claim fails.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the HRIES scale, the 16-item instrument used to measure perceived anthropomorphism of each chatbot."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes the precedent of connecting usage motivations to media dependency, applied here to Facebook."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Extends media dependency thinking to AI companionship, the direct theoretical background for dependency on chatbots."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the grounded-theory method used to code interview responses and derive the three interfering factors."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Sets the Cronbach's alpha acceptance threshold used to justify scale reliability."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Shows media dependency theory applied to an AI-driven app, grounding the model's satisfaction-to-attitude-to-behavior chain."}],"review_version":1}