{"id":"2f8102ef-aa12-444b-af38-3dc356d0079d","arxiv_id":"2412.19976","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Adding name and background to a donation chatbot increased perceived humanness but did not increase donations; a plain chatbot with logical appeals was viewed most favorably.","lead":"A lab experiment with 76 students tested whether giving a donation chatbot a name, avatar, and backstory increases donations. It did not, and a non-personified chatbot using logical appeals was rated most favorably.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Reported coefficient in §4.1.2 contradicts the paper's central claim: personification is declared to increase mindless anthropomorphism, but the reported β = -1.69 indicates the opposite under the stated scale.","rationale":"I read the paper in good faith and initially focused on the ceiling effect for willingness to donate (§3.5.1: 81.58% donated the full $10) and the reader's point about mediation ordering (§3.5.1 measures donation before the post-interaction perception questionnaire in §3.1). Both are real limitations. However, the most load-bearing problem is internal and more basic: the reported regression coefficient for personification on mindless anthropomorphism contradicts the paper's own conclusion. The scale direction is unambiguous in §3.5.3, and the interpretation of the chatbot-perceptions result in §4.1.3 fixes the dummy coding. A negative coefficient for the focal condition on a higher-is-more-anthropomorphic scale cannot support H1b. This is not a matter of external consensus or statistical power; it is an inconsistency within the reported results. If the numbers are correct, the paper's theoretical mechanism (personification → perceived anthropomorphism) is not merely unsupported but contradicted. If the sign is a typo, the analysis must be rerun and every downstream claim re-evaluated. I therefore recommend a conditional verdict: acceptance requires the authors to resolve this inconsistency with a reanalysis of the raw data. The reader's mediation-ordering concern is also valid and should be addressed, but the sign contradiction is more fundamental. I give partial agreement with the reader because their identified weak assumption is real yet not the single most load-bearing issue.","tokens_in":11199,"tokens_out":5845,"duration_ms":58379,"concrete_test":"Obtain the raw data (or a full covariance/correlation matrix) for the variables in §4.1.2, including personification dummy coding, persuasion strategy, mindless anthropomorphism, mindful anthropomorphism, chatbot perceptions, familiarity, trait empathy, and WTD. Re-estimate the OLS model for mindless anthropomorphism. If the personification coefficient remains -1.69 (or negative) with personified=1 and higher scores indicating more anthropomorphism, then H1b is rejected and the central claim is unsupported. If the negative sign is a typographical error, the corrected coefficient must be positive and significant; all subsequent mediation results and the abstract's first clause must be recomputed before the paper can be relied upon.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The paper's central claim requires that personification evokes perceived anthropomorphism. In §4.1.2, the model predicting mindless anthropomorphism reports: 'Chatbot personification, β = -1.69, t(70) = -2.32, p = .02.' The measure in §3.5.3 is scored so that higher values indicate more anthropomorphism (attractive, exciting, pleasant, interesting, likable, sociable, friendly, personal). The same analysis in §4.1.3 interprets β = -1.36 for personification on chatbot perceptions as 'personification actually led to less favorable chatbot perceptions,' confirming that personified is the focal coding. Under this coding, β = -1.69 means personification *decreases* mindless anthropomorphism, the opposite of H1b. The paragraph nevertheless says 'H1b was supported,' and the abstract and discussion claim personified chatbots evoke perceived anthropomorphism. Either the sign, the variable coding, or the conclusion is erroneous. If the coefficient is correct, the first half of the central claim collapses; the mediation results that follow rest on a rejected manipulation effect. The reader's temporal-ordering concern about mediation is secondary; this sign inconsistency is more load-bearing because it undermines the hypothesized X→M path before mediation is even tested.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a between-subjects experiment (N = 76) crossing chatbot personification (name and background information vs. none) with persuasion strategy (emotional vs. logical) in a donation context, and tests whether perceived anthropomorphism (mindful and mindless) and chatbot perceptions mediate effects on willingness to donate (WTD) using PROCESS Model 7. The headline claims are that personified chatbots increase mindless anthropomorphism but do not increase WTD, and that personification can lead to negative chatbot attitudes, especially when paired with a logical appeal. The paper also reports moderated mediation results and discusses implications for AI disclosure regulations.","tokens_in":11486,"tokens_out":3494,"duration_ms":34786,"significance":"If the reported results were internally consistent, the study would be a useful counterpoint to CASA-based expectations that anthropomorphic cues improve chatbot outcomes, and it would extend congruency arguments to donation interactions. The paper is transparent about several limitations, including the small homogeneous sample and the ceiling effect in WTD, and it makes its coefficients and confidence intervals available in the text. However, the central claim is undermined by a sign inconsistency in the main predictor's effect on the primary mediator, and the mediation analysis has a temporal-order problem. These issues prevent the paper from supporting its abstract and discussion claims as written, though the underlying questions remain of interest to the HCI community.","major_comments":[{"comment":"The reported coefficient for chatbot personification on mindless anthropomorphism is β = -1.69, t(70) = -2.32, p = .02. Given that the mindless anthropomorphism scale is scored so that higher values indicate more anthropomorphism (§3.5.3), and given that §4.1.3 interprets a negative coefficient as showing that personification 'led to less favorable chatbot perceptions,' this negative sign means personification decreased mindless anthropomorphism, the opposite of H1b. The text nevertheless states 'H1b was supported.' This contradiction directly undermines the abstract's claim that a personified chatbot evokes perceived anthropomorphism and removes the X→M path required for the subsequent mediation claims. The authors must clarify the variable coding or correct the conclusion; if the coefficient is correct, the first half of the central claim collapses.","section":"§4.1.2"},{"comment":"WTD was elicited during the chatbot conversation (before the end of the interaction), while the mediator measures (mindful and mindless anthropomorphism, chatbot perceptions) were collected in the post-interaction questionnaire. The mediation model in Figure 2 treats the mediators as intervening between the manipulation and the outcome, but the measurement order is reversed. Consequently, the estimated indirect effects cannot support the proposed causal direction from personification through perceptions to WTD. The authors should either report the analysis as correlational with explicit caveats, or redesign the measurement order in future work.","section":"§3.5.1 and §3.1"},{"comment":"The WTD variable exhibits a severe ceiling effect: 81.58% of participants donated the full $10, and the mean is 8.66 (SD = 3.01) on a 0–10 scale. This near-dichotomous distribution severely limits variance and power. All bootstrap confidence intervals for indirect effects include zero, so the null mediation results may reflect measurement insensitivity rather than a true absence of effect. The limitations section acknowledges the skew but does not assess its impact; the authors should report alternative analyses (e.g., logistic or ordinal regression on donation amount, or a sensitivity analysis excluding or transforming the ceiling cases).","section":"§3.5.1 and §7"},{"comment":"Ten participants were excluded for failing the manipulation check, but no sensitivity analysis is reported. If exclusion rates differ by condition or correlate with outcomes, the reported estimates could be biased. The authors should report the number excluded per condition and rerun the key models with the full sample to establish robustness.","section":"§3.2 and §4"},{"comment":"The mindless anthropomorphism scale (attractive, exciting, pleasant, interesting, likable, sociable, friendly, personal) shares items (pleasant, interesting) with the chatbot experience component used in the aggregated chatbot perceptions measure. This construct overlap calls into question the discriminant validity of the two mediators and may inflate their intercorrelation, complicating the interpretation of the mediation paths. The authors should address this overlap, either by reporting a CFA or by re-analyzing with non-overlapping items.","section":"§3.5.2 and §3.5.3"}],"minor_comments":[{"comment":"The participant age is reported as 'M = 21,06; SD = 2.39,' which appears to be a typo for 'M = 21.06.'","section":"§3.2"},{"comment":"The aggregated chatbot perceptions mean is reported as 'M = 4.81, SD = .13'; the standard deviation of .13 seems implausibly small relative to the component scales (SD = 1.04 and 1.29) and is likely a typo for SD ≈ 1.13.","section":"§3.5.2"},{"comment":"The interaction coefficient is reported as 'β = .979'; for readability and consistency with the other coefficients, it should be formatted as 'β = 0.98.'","section":"§4.1.3"},{"comment":"The Figure 2 description in the appendix is detailed and helpful, but it would be clearer if the significance labels (S, NS, S*) were defined in the main text near the figure rather than only in the appendix.","section":"Appendix"}],"recommendation":"major_revision","confidential_remarks":"The sign inconsistency in §4.1.2 is the most serious issue: if the reported β = -1.69 is correct, the abstract's claim that personification evokes anthropomorphism is false and the mediation framework collapses. If it is a reporting error, the manuscript needs a full reanalysis and a revised narrative. The temporal-order problem with WTD measured before the mediators is also a fundamental design issue that cannot be fixed by rewriting alone. The paper is within scope for an HCI empirical study, but the current version does not support its central claims."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague, quick take on arXiv:2412.19976. The study is a straightforward 2x2 lab experiment (personified vs. non-personified chatbot, emotional vs. logical appeal) on donation willingness from a small student sample. The novel bit is the interaction between personification and persuasion strategy, and the finding that a non-personified chatbot with a logical appeal is rated most favorably. That is a genuinely useful design insight for nonprofits, and the paper connects it sensibly to bot-disclosure regulation. The writing is clear and the literature review is honest enough.\n\nBut there is a load-bearing problem. Section 4.1.2 reports that personification significantly predicts mindless anthropomorphism with β = -1.69, and then says \"H1b was supported.\" The mindless anthropomorphism scale is scored so higher values mean more anthropomorphic (attractive, sociable, personal, etc.), and the same coefficient in Section 4.1.3 is interpreted as personification making things worse. So the sign says personification decreases mindless anthropomorphism, not increases it. Either the coding is wrong, or the conclusion is. The abstract and discussion lean on personification evoking perceived anthropomorphism, and the mediation model in Section 4.2 rests on that path. If the coefficient is right, the first half of the central claim collapses.\n\nThere are smaller issues. The WTD measure has a severe ceiling effect: 81.6% donated the full $10, which makes the null direct effect nearly uninformative. The mediators were measured after the donation question, so the mediation analysis violates temporal precedence. Ten participants were dropped for failing the manipulation check, with no sensitivity analysis. No data or code are shared, so the sign issue cannot be resolved by inspection.\n\nWhat the paper does well is the interaction framing and the practical question it asks. But as written, the main effect it reports is internally inconsistent, and the mediation story is not identifiable from the design.\n\nFor peer review: I would not desk-reject this outright, because a reanalysis with a corrected sign and a proper donation measure could yield something worth publishing. But a serious referee would need to send it back for major revision, and the authors should be required to release data and preregister a follow-up. As it stands, I would not cite the findings without first checking the sign. Worth a reading group as a cautionary example of mediation timing.\n\nRecommendation: send to peer review only if the editor is willing to demand a full reanalysis; otherwise desk-reject on the internal contradiction.","headline":"A clean 2x2 study that would be worth a look, but the central H1b claim is contradicted by the paper's own reported coefficient, and the mediation design breaks temporal order.","tokens_in":11932,"tokens_out":2069,"would_cite":false,"duration_ms":23233,"reading_group":"maybe","serious_thinker":"no","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Giving a donation chatbot a name and a backstory makes people see it as more human but does not make them more willing to donate, and can even sour their attitudes toward it.","keywords":["chatbot anthropomorphism","willingness to donate","personification","persuasion strategy","mindless anthropomorphism","human-chatbot interaction","nonprofit fundraising","AI disclosure"],"falsifier":"A replication that measures perceived anthropomorphism and chatbot perceptions before the donation request—rather than after the full conversation—would settle whether the proposed causal chain holds. If the indirect effects of personification through perceptions become significant, or change sign, when perceptions are measured pre-donation, the current mediation conclusions would be refuted. A second decisive check is to replace the $10 capped endpoint with an open donation amount and recruit a larger, more donation-representative sample; if the direct and indirect effects remain null under those conditions, the 'no benefit from personification' claim is much stronger.","tokens_in":11034,"feed_emoji":"🤖","tokens_out":5835,"duration_ms":57992,"temperature":0.7,"pith_summary":"This paper tests a common assumption behind charity chatbot design: that making a bot feel more human—by giving it a name, an avatar, and a personal story—will make people more willing to donate. In a 2x2 experiment (personified vs. non-personified chatbot, emotional vs. logical persuasion) with 76 participants, the personified bot did increase perceived anthropomorphism of the unconscious, 'mindless' kind, but it did not increase the amount people chose to donate. Instead, personification produced significantly less favorable overall chatbot perceptions, and this negative effect appeared when the bot used a logical, statistics-based appeal. The authors conclude that anthropomorphic cues are not a reliable route to donation behavior, and that in the donation context a plainly non-human bot giving logical reasons may be the safer design.","feed_headline":"Chatbot name and story backfire on donation pages","feed_subtitle":"Plain bots with statistics won the most favorable reactions in a 76-person giving experiment.","key_machinery":"The analytical machinery is a moderated mediation model estimated with PROCESS Model 7, in which chatbot personification is the predictor, willingness to donate is the outcome, perceived mindful anthropomorphism, perceived mindless anthropomorphism, and chatbot perceptions are parallel mediators, and persuasion strategy moderates the path from personification to the mediators. The conceptual load-bearing distinction comes from prior work separating mindful anthropomorphism (deliberate attribution of human-likeness to a non-human) from mindless anthropomorphism (the automatic, reflexive version). The model's results show that personification moves only the mindless pathway, that this pathway does not connect to donation behavior, and that the moderation pattern—personification hurting perceptions under logical appeals but not under emotional appeals—supports a consistency explanation of the findings.","core_discovery":"The study's central claim is that commonly used anthropomorphic cues—a name, an avatar, and a personal background narrative—trigger mindless anthropomorphism (the reflexive sense that the agent is human-like) without triggering mindful anthropomorphism, and this perceptual shift does not carry over into willingness to donate. The direct effect of personification on donation amount was not significant, and the personified chatbot actually received significantly worse chatbot perceptions than the non-personified version, opposite to the authors' hypothesis. Moderation analyses show this negative effect was strongest in the logical-persuasion condition, while in the emotional-persuasion condition the personified and non-personified bots were perceived similarly. Read together with prior work on cue consistency, the authors argue that the fit between the chatbot's identity and its persuasive style matters: a machine-like agent that speaks in statistics feels more coherent than a humanized agent that does the same.","pith_inferences":["If the consistency explanation generalizes, the same identity-persuasion mismatch should predict attitudes in other contexts: a personified assistant using dry logic would be evaluated poorly, while a clearly mechanical bot using emotional language may also seem inconsistent—a testable crossover interaction.","Because donation amount was elicited during the conversation and perception ratings came after it, the mediation estimates depend on the assumption that post-conversation ratings reflect pre-donation perceptions; measuring perceptions before the donation request in a follow-up would directly test that assumption.","The results suggest a possible negative mechanism beyond congruency: a humanized bot citing statistics may be seen as manipulative, while a non-human bot doing the same is seen as objective; measuring perceived sincerity or manipulativeness would separate these accounts.","The study's sample is small and homogeneous, so the null indirect effects are weak evidence of absence; a sufficiently powered replication with a continuous donation measure could reveal whether the absence of mediation is real."],"forward_implications":["Charity chatbot designers should not assume that human-like features increase donations; in this study, a name and backstory made attitudes worse rather than better.","A non-personified chatbot using logical, data-driven appeals produced the most favorable chatbot perceptions of the four conditions, suggesting this combination as a baseline design for donation contexts.","The study separates two routes to anthropomorphism: cues can trigger the mindless route without the mindful route, so evaluations and donation behavior should be measured separately.","The null mediation results imply that even when anthropomorphism is successfully evoked, it does not automatically act as a bridge to desired behavioral outcomes such as giving.","Because 81.6% of participants donated the full $10, the donation measure was near its ceiling; the paper's own limitation note implies the null effects on donation amount may be partly an artifact of this skew."],"supporting_citations":[{"why":"Provides the prior evidence that anthropomorphic design cues like a human name raise perceived anthropomorphism and improve company attitudes—the starting point the experiment extends to donations.","marker":"[2]"},{"why":"Supplies the mindful/mindless anthropomorphism distinction that the study uses as its two mediators.","marker":"[25]"},{"why":"Previous chatbot donation study showing a bot with a human name increased perceived humanness and willingness to donate, the direct comparison the current results complicate.","marker":"[40]"},{"why":"Establishes the consistency preference for agent cues that the authors use to explain why persona-persuasion mismatches hurt perceptions.","marker":"[15]"},{"why":"Shows human-like chatbot language can impair charitable giving, supporting the negative-effect finding in a donation context.","marker":"[49]"},{"why":"Provides the PROCESS Model 7 moderated mediation method used to estimate direct and indirect effects.","marker":"[18]"}],"fun_headline_variants":["Personified chatbot fails to boost donations, hurts perception","Donors prefer logical plain chatbots over human-like ones","Anthropomorphic chatbot cues lower donation willingness"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that the perception ratings collected after the chatbot conversation reflect the same perceptions participants held when they chose their donation amount during the conversation; if making the donation choice changed their later ratings, the mediation analysis cannot establish the proposed causal order from personification through perceptions to donation behavior.","fun_headline_variants_meta":{"raw":{"variants":["Personified chatbot fails to boost donations, hurts perception","Donors prefer logical plain chatbots over human-like ones","Anthropomorphic chatbot cues lower donation willingness"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000657,"raw_usage":{"total_tokens":2978,"prompt_tokens":889,"completion_tokens":2089,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":505,"completion_tokens_details":{"reasoning_tokens":2041}},"tokens_in":505,"tokens_out":2089,"duration_ms":15911,"temperature":1.0,"reasoning_tokens":2041,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T23:42:48.944390+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A replication that measures perceived anthropomorphism and chatbot perceptions before the donation request—rather than after the full conversation—would settle whether the proposed causal chain holds. If the indirect effects of personification through perceptions become significant, or change sign, when perceptions are measured pre-donation, the current mediation conclusions would be refuted. A second decisive check is to replace the $10 capped endpoint with an open donation amount and recruit a larger, more donation-representative sample; if the direct and indirect effects remain null under those conditions, the 'no benefit from personification' claim is much stronger.","supporting_citations":[{"cited_title":"2020, pp","cited_arxiv_id":null,"evidence_quote":"Supplies the mindful/mindless anthropomorphism distinction that the study uses as its two mediators."},{"cited_title":"Effects of visual and linguistic anthropomorphic cues on social perception, self-awareness, and information disclosure in a health website","cited_arxiv_id":null,"evidence_quote":"Previous chatbot donation study showing a bot with a human name increased perceived humanness and willingness to donate, the direct comparison the current results complicate."},{"cited_title":"How human–chatbot interaction impairs charitable giving: the role of moral judgment","cited_arxiv_id":null,"evidence_quote":"Shows human-like chatbot language can impair charitable giving, supporting the negative-effect finding in a donation context."}],"review_version":1}